CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ParlAI /blended_skill_talk Dataset Card for "blended_skill_talk" Dataset Summary A dataset of 7k conversations explicitly designed to exhibit multiple conversation modes: displaying personality, having empathy, and demonstrating knowledge. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 38.11 MB Size of the generated dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ParlAI/blended_skill_talk.text1K<n<10K75 likes6.9k downloads3y agoHugging Face02llm-blender /Unified-FeedbackCollections of pairwise feedback datasets. openai/summarize_from_feedback openai/webgpt_comparisons Dahoas/instruct-synthetic-prompt-responses Anthropic/hh-rlhf lmsys/chatbot_arena_conversations openbmb/UltraFeedback argilla/ultrafeedback-binarized-preferences-cleaned berkeley-nest/Nectar Codes to reproduce the dataset: jdf-prog/UnifiedFeedback Dataset formats { "id": "...", "conv_A": [ { "role": "user", "content": "...", }, { "role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/Unified-Feedback.tabular1M<n<10M18 likes4.6k downloads2y agoHugging Face03BleachNick /UltraEdit_500k Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/BleachNick/UltraEdit_500k.imagetext-to-image100K<n<1M15 likes1.9k downloads2y agoHugging Face04BleachNick /UltraEdit_Region_Based_100k Bibtex citation @misc{zhao2024ultraeditinstructionbasedfinegrainedimage, title={UltraEdit: Instruction-based Fine-Grained Image Editing at Scale}, author={Haozhe Zhao and Xiaojian Ma and Liang Chen and Shuzheng Si and Rujie Wu and Kaikai An and Peiyu Yu and Minjia Zhang and Qing Li and Baobao Chang}, year={2024}, eprint={2407.05282}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2407.05282}, } imagetext-to-image100K<n<1M11 likes891 downloads2y agoHugging Face05MaxRondelli /BlenderRAG BlenderRAG Dataset A dataset for 3D scene and object generation research. Each sample pairs a Blender Python script that procedurally generates a 3D object with a rendered preview image and a natural-language description. Dataset Summary The dataset is organized into two top-level scenes — indoor and outdoor — each containing a collection of objects. Every object is represented by three aligned modalities: File Modality Purpose code_n.py Python (Blender API)… See the full description on the dataset page: https://huggingface.co/datasets/MaxRondelli/BlenderRAG.imagen<1K4 likes881 downloads5mo agoHugging Face06BleachNick /MIC_sampledtext100K<n<1M3 likes724 downloads3y agoHugging Face07yanglll /blendtext1M<n<10M0 likes528 downloads2y agoHugging Face08Incomple /BLEnD-Vis BLEnD-Vis BLEnD-Vis is a benchmark for evaluating vision-language models (VLMs) on culturally grounded multiple-choice questions, including a text-only setting and a visual setting with generated images. Paper: https://arxiv.org/abs/2510.11178 Dataset repo: https://huggingface.co/datasets/Incomple/BLEnD-Vis Code: https://github.com/Social-AI-Studio/BLEnD-Vis Source BLEnD-Vis is derived from the BLEnD dataset on Hugging Face (nayeon212/BLEnD). What is in… See the full description on the dataset page: https://huggingface.co/datasets/Incomple/BLEnD-Vis.imagevisual-question-answering10K<n<100K0 likes517 downloads8mo agoHugging Face09RoboCOIN /AIRBOT_MMK2_store_beauty_blender_and_building_blocksgated AIRBOT_MMK2_store_beauty_blender_and_building_blocks 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick 📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_store_beauty_blender_and_building_blocks.tabularrobotics1K<n<10K0 likes485 downloads9mo agoHugging Face10RoboCOIN /AIRBOT_MMK2_pour_out_the_beauty_blendergated AIRBOT_MMK2_pour_out_the_beauty_blender 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_pour_out_the_beauty_blender.tabularrobotics10K<n<100K0 likes438 downloads9mo agoHugging Face11autobio-bench /thermal_mixer-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 79278, "total_tasks": 100, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/thermal_mixer-blender.tabularrobotics10K<n<100K0 likes360 downloads1y agoHugging Face12autobio-bench /thermal_cycler_close-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 106292, "total_tasks": 1, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/thermal_cycler_close-blender.tabularrobotics100K<n<1M0 likes300 downloads1y agoHugging Face13autobio-bench /pipette-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 71014, "total_tasks": 1, "total_videos": 300, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/pipette-blender.tabularrobotics10K<n<100K0 likes275 downloads1y agoHugging Face14autobio-bench /thermal_cycler_open-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 84994, "total_tasks": 1, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/thermal_cycler_open-blender.tabularrobotics10K<n<100K0 likes271 downloads1y agoHugging Face15autobio-bench /insert-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 55127, "total_tasks": 10, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/insert-blender.tabularrobotics10K<n<100K0 likes231 downloads1y agoHugging Face16Lots-of-LoRAs /task1418_bless_semantic_relation_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1418_bless_semantic_relation_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1418_bless_semantic_relation_classification.texttext-generation1K<n<10K0 likes202 downloads2y agoHugging Face17Blebbyblub /Indspeech-Augmented-Dataset-10000-20000audio10K<n<100K0 likes154 downloads1y agoHugging Face18autobio-bench /insert_centrifuge_5430-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 54329, "total_tasks": 1, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/insert_centrifuge_5430-blender.tabularrobotics10K<n<100K0 likes151 downloads1y agoHugging Face19ThomasTheMaker /BlenderCAD2imagen<1K0 likes133 downloads1y agoHugging Face20autobio-bench /screw_tighten-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 154830, "total_tasks": 1, "total_videos": 300, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/screw_tighten-blender.tabularrobotics100K<n<1M0 likes119 downloads1y agoHugging Face21yapeichang /BLEUBERI-Tulu3-50k[Paper] [HF Collection] [Code] Authors: Yapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh, Chuan Li, Chris Tanner, Mohit Iyyer Contact: yapeic@umd.edu TLDR > We extend RLVR beyond easily verifiable domains like math and code to the more open-ended setting of general instruction following. Surprisingly, we find that BLEU—a simple n-gram matching metric—when paired with high-quality references from strong LLMs, achieves human agreement comparable to 8B and 27B reward models on Chatbot… See the full description on the dataset page: https://huggingface.co/datasets/yapeichang/BLEUBERI-Tulu3-50k.texttext-generation10K<n<100K2 likes83 downloads1y agoHugging Face22hugo0076 /CIFAR100-Blended-20pct-Backdoor-ExclNaturalimage10K<n<100K0 likes81 downloads11mo agoHugging Face23Lots-of-LoRAs /task1582_bless_hypernym_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1582_bless_hypernym_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1582_bless_hypernym_generation.texttext-generationn<1K0 likes78 downloads2y agoHugging Face24autobio-bench /pickup-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 49959, "total_tasks": 1, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/pickup-blender.tabularrobotics10K<n<100K0 likes78 downloads1y agoHugging Face25autobio-bench /screw_loose-blenderThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 136834, "total_tasks": 1, "total_videos": 300, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/autobio-bench/screw_loose-blender.tabularrobotics100K<n<1M0 likes74 downloads1y agoHugging Face26Lots-of-LoRAs /task1583_bless_meronym_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1583_bless_meronym_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1583_bless_meronym_classification.texttext-generation1K<n<10K0 likes70 downloads2y agoHugging Face27bleondubos /red-team-unified-datasettext100K<n<1M0 likes68 downloads8d agoHugging Face28BleachNick /CraftBenchimagen<1K1 likes65 downloads4mo agoHugging Face29CooperBench /cooperdata-v3-midtrain-blend CooperData v3 — Midtraining Blend (Qwen3.5-9B cooperative SWE agents) All-token midtraining mixture that bridges Qwen/Qwen3.5-9B (instruct) toward the cooperative multi-agent SWE-coding SFT distribution. One document per row (text, tagged by source) — NOT packed — so trl.SFTTrainer(packing=False) tokenizes per-doc and the Gated-DeltaNet recurrence stays per-document. ~390M tokens. Composition source tokens share role web 210.0M 54% general math 55.0M… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/cooperdata-v3-midtrain-blend.texttext-generation100K<n<1M0 likes58 downloads3mo agoHugging Face30NagaYu /bleep-spans Bleep spans — synthetic sensitive-speech regions with frame-accurate labels Where sensitive information is spoken, and what kind it is — never what was said. Every recording is synthetic. No real telephone call, clinical recording, or any other real speech was used, recorded, or derived from at any stage. 🤗 Model: NagaYu/bleep-0.09b 🎛️ Demo: NagaYu/bleep What a row contains utt_id, voice_key, condition, duration, subsets, and three parallel arrays —… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/bleep-spans.audioaudio-classification1K<n<10K0 likes57 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.