CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openai /openai_humaneval Dataset Card for OpenAI HumanEval Dataset Summary The HumanEval dataset released by OpenAI includes 164 programming problems with a function sig- nature, docstring, body, and several unit tests. They were handwritten to ensure not to be included in the training set of code generation models. Supported Tasks and Leaderboards Languages The programming problems are written in Python and contain English natural text in comments and docstrings.… See the full description on the dataset page: https://huggingface.co/datasets/openai/openai_humaneval.textn<1K404 likes265k downloads3y agoHugging Face02evalplus /humanevalplustextn<1K23 likes31k downloads2y agoHugging Face03lerobot /aloha_sim_transfer_cube_humanThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 20000, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_human.tabularrobotics10K<n<100K16 likes19k downloads4mo agoHugging Face04bigcode /humanevalpack Dataset Card for HumanEvalPack Dataset Summary HumanEvalPack is an extension of OpenAI's HumanEval to cover 6 total languages across 3 tasks. The Python split is exactly the same as OpenAI's Python HumanEval. The other splits are translated by humans (similar to HumanEval-X but with additional cleaning, see here). Refer to the OctoPack paper for more details. Languages: Python, JavaScript, Java, Go, C++, Rust OctoPack🐙🎒: Data CommitPack 4TB of GitHub commits… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/humanevalpack.textn<1K92 likes11k downloads1y agoHugging Face05meoconxinhxan /Medical-Eval-HumanityLastExamtextn<1K1 likes11k downloads1y agoHugging Face06lerobot /aloha_sim_insertion_humanThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 25000, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_human.tabularrobotics10K<n<100K13 likes9.7k downloads4mo agoHugging Face07iNeil77 /HumanEval-XLThis dataset contains a viewer-friendly version of the dataset at FloatAI/HumanEval-XL. It is made available separately for the convenience of the vllm-code-harness package. texttext-generation10K<n<100K0 likes9.7k downloads2y agoHugging Face08lerobot /aloha_sim_insertion_human_imageThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 25000, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null, "features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_human_image.imagerobotics10K<n<100K1 likes4.4k downloads4mo agoHugging Face09vidore /esg_reports_human_labeled_v2 Vidore Benchmark 2 - ESG Human Labeled This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry. Dataset Summary Each query is in english. This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.imagedocument-question-answering1K<n<10K1 likes4.1k downloads1y agoHugging Face10lerobot /aloha_sim_transfer_cube_human_imageThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 20000, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null, "features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_human_image.imagerobotics10K<n<100K2 likes4.1k downloads4mo agoHugging Face11humanify /si_for_sdaudio10K<n<100K0 likes3.9k downloads2mo agoHugging Face12HumanBehaviorAtlas /human_behavior_atlas Human Behavior Atlas A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features. This dataset was used to train OmniSapiens, a foundation model for social behavior processing. Papers: Human Behavior Atlas: Benchmarking Unified Psychological and… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas.textvideo-classification100K<n<1M3 likes3.3k downloads4mo agoHugging Face13lmarena-ai /arena-human-preference-140k Overview This dataset contains user votes collected in the text-only category. Each row represents a single vote judging two models (model_a and model_b) on a user conversation, along with the full conversation history and metadata. Key fields include: id: Unique feedback ID of each vote/row. evaluation_session_id: Unique ID of each evaluation session, which can contain multiple separate votes/evaluations. evaluation_order: Evaluation order of the current vote. winner: Battle… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-140k.text100K<n<1M56 likes3.1k downloads1y agoHugging Face14DynamicIntelligence /humanoid-robots-training-dataset Dynamic Intelligence — Humanoid Robot Training Dataset A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors. The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.tabularrobotics10K<n<100K0 likes2.8k downloads6mo agoHugging Face15MBZUAI /human_translated_arabic_mmlutext10K<n<100K4 likes2.2k downloads2y agoHugging Face16lmsys /mt_bench_human_judgments Content This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions. The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper. Agreement Calculation This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/mt_bench_human_judgments.tabularquestion-answering1K<n<10K148 likes2k downloads3y agoHugging Face17GalaxyGeneralRobotics /HumanTracker Dataset Card for HumanTracker Project page · Paper · Code HumanTracker is a humanoid motion-tracking benchmark. This release contains two complementary subsets: motions/ — the evaluation test split: retargeted 29-DoF reference trajectories, grouped into four motion families. preference_pair/ — 6,000 human preference pairs, each stored with the two tracker rollouts that were compared and the source-motion clip they track. The evaluation harness and HumanScore reward model live… See the full description on the dataset page: https://huggingface.co/datasets/GalaxyGeneralRobotics/HumanTracker.tabularrobotics1K<n<10K6 likes1.9k downloads29d agoHugging Face18Rapidata /text-2-image-Rich-Human-Feedback Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.imagetext-to-image10K<n<100K37 likes1.8k downloads2y agoHugging Face19skyzos /Humanoid-Everyday-G1tabular1M<n<10M0 likes1.6k downloads9mo agoHugging Face20lerobot /robocasa_target_human_unifiedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "robocasa", "total_episodes": 25307, "total_frames": 14957899, "total_tasks": 50, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:25307" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/robocasa_target_human_unified.tabularrobotics10M<n<100M4 likes1.6k downloads5mo agoHugging Face21humanify /Env-TTS-Clean Env-TTS-Clean Environment-aware text-to-speech training corpus (clean release). Each row pairs four short 24 kHz mono FLAC clips with aligned transcripts: an environment sample (different speaker, same acoustic scene), a speaker reference (same speaker as the target utterance), a speaker-enhanced copy of the reference (MossFormer2 enhancement — or, for the DDS source, the real clean-studio recording of the speaker reference), the target speech to synthesise, so a model can… See the full description on the dataset page: https://huggingface.co/datasets/humanify/Env-TTS-Clean.audio100K<n<1M0 likes1.5k downloads2mo agoHugging Face22ellamind /humaneval-multilingualtextn<1K0 likes1.4k downloads7mo agoHugging Face23lmarena-ai /arena-human-preference-100k Overview This dataset contains leaderboard conversation data collected between June 2024 and August 2024. It includes English human preference evaluations used to develop Arena Explorer. Additionally, we provide an embedding file, which contains precomputed embeddings for the English conversations. These embeddings are used in the topic modeling pipeline to categorize and analyze these conversations. For a detailed exploration of the dataset and analysis methods, refer to the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-100k.tabular100K<n<1M49 likes1.3k downloads2y agoHugging Face24USC-PSI-Lab /Humanoid-Everyday-G1tabular1M<n<10M0 likes1.2k downloads10mo agoHugging Face25BryanW /HumanEdit Dataset Card for HumanEdit Paper (CVPR 2025 AI for Content Creation (AI4CC) Workshop) Usage from datasets import load_dataset from PIL import Image # Load the dataset ds = load_dataset("BryanW/HumanEdit") # Print the total number of samples and show the first sample print(f"Total number of samples: {len(ds['train'])}") print("First sample in the dataset:", ds['train'][0]) # Retrieve the first sample's data data_dict = ds['train'][0] # Save the input image (INPUT_IMG)… See the full description on the dataset page: https://huggingface.co/datasets/BryanW/HumanEdit.imagetext-to-image1K<n<10K25 likes1.2k downloads1y agoHugging Face26Rapidata /text-2-video-human-preferences Rapidata Video Generation Preference Dataset This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set: Sora Hunyouan Pika 2.0 Runway ML Alpha Luma Ray 2 Explore our latest model rankings on our website. If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.imagetext-to-video1K<n<10K21 likes1.1k downloads2y agoHugging Face27lmarena-ai /PPE-Human-Preference-V1 Overview This contains the human preference evaluation set for Preference Proxy Evaluations. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward Models for RLHF}, author={Evan Frick and Tianle Li and… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-Human-Preference-V1.tabular10K<n<100K10 likes1.1k downloads2y agoHugging Face28pepijn223 /robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "observation.images.robot0_agentview_left": { "dtype": "video", "shape": [ 256, 256, 3 ], "names": [ "height", "width", "channel" ], "video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.tabularrobotics10M<n<100M2 likes1.1k downloads3mo agoHugging Face29Rapidata /text-2-video-human-preferences-wan2.1 Rapidata Video Generation Alibaba Wan2.1 Human Preference If you get value from this dataset and would like to see more in the future, please consider liking it. This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Overview In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.imagevideo-classificationn<1K20 likes1.1k downloads2y agoHugging Face30Rapidata /human-coherence-preferences-images Rapidata Image Generation Coherence Dataset This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.imagetext-to-image10K<n<100K14 likes991 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.