datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UltraEdit
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/BleachNick/UltraEdit.BlendedMVS_processedBlenderLore
The target scope is 22,360 video-associated Blender project instances, not 22,360 distinct tutorial videos. Uploads are in progress, so the currently published files may be a subset of this target. The 44 biomedical project instances and one software-bundled Dome template are excluded.
Data Structure
The dataset is organized as a collection of sample-level directories under assets/. Each directory corresponds to one Blender creation task and follows the structure below:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/BlenderLore.bleach
Bangumi Image Base of Bleach
This is the image base of bangumi Bleach, we detected 181 characters, 30903 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters' preview:… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/bleach.blended_skill_talk
Dataset Card for "blended_skill_talk"
Dataset Summary
A dataset of 7k conversations explicitly designed to exhibit multiple conversation modes: displaying personality, having empathy, and demonstrating knowledge.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 38.11 MB
Size of the generated dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ParlAI/blended_skill_talk.Unified-FeedbackCollections of pairwise feedback datasets.
openai/summarize_from_feedback
openai/webgpt_comparisons
Dahoas/instruct-synthetic-prompt-responses
Anthropic/hh-rlhf
lmsys/chatbot_arena_conversations
openbmb/UltraFeedback
argilla/ultrafeedback-binarized-preferences-cleaned
berkeley-nest/Nectar
Codes to reproduce the dataset: jdf-prog/UnifiedFeedback
Dataset formats
{
"id": "...",
"conv_A": [
{
"role": "user",
"content": "...",
},
{
"role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/Unified-Feedback.blendsql-test-dbsHoyoverse_Character_ModelsUltraEdit_500k
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/BleachNick/UltraEdit_500k.cua-blenderBLEnD
BLEnD
This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track).
24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.Nemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.BlenderBench
BlenderBench Dataset
Dataset Description
BlenderBench is a comprehensive benchmark dataset for evaluating models on 3D scene editing tasks in Blender. The dataset challenges agents to understand visual differences between initial and target scenes, then generate appropriate Blender Python code to transform the initial scene to match the target.
Key Features
27 instances across 3 difficulty levels
Multi-modal: Combines visual… See the full description on the dataset page: https://huggingface.co/datasets/DietCoke4671/BlenderBench.NOAA-PIFSC-ESD-CORAL-Bleaching-Dataset
Dataset Card for NOAA-ESD-CORAL-Bleaching Classification Dataset v1
Overview
For the development of machine learning models to classify coral health, specifically identifying healthy hard coral (CORAL) and bleached hard coral (CORAL_BL).This dataset contains underwater imagery collected by NOAA's Ecosystem Sciences Division (ESD) and other benthic surveys.
Labels
Label
Name
Functional Group
CORAL
Healthy Hard Coral
Hard Coral
CORAL_BL… See the full description on the dataset page: https://huggingface.co/datasets/De129472/NOAA-PIFSC-ESD-CORAL-Bleaching-Dataset.UltraEdit_Region_Based_100k
Bibtex citation
@misc{zhao2024ultraeditinstructionbasedfinegrainedimage,
title={UltraEdit: Instruction-based Fine-Grained Image Editing at Scale},
author={Haozhe Zhao and Xiaojian Ma and Liang Chen and Shuzheng Si and Rujie Wu and Kaikai An and Peiyu Yu and Minjia Zhang and Qing Li and Baobao Chang},
year={2024},
eprint={2407.05282},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2407.05282},
}
BlenderRAG
BlenderRAG Dataset
A dataset for 3D scene and object generation research. Each sample pairs a Blender Python script that procedurally generates a 3D object with a rendered preview image and a natural-language description.
Dataset Summary
The dataset is organized into two top-level scenes — indoor and outdoor — each containing a collection of objects. Every object is represented by three aligned modalities:
File
Modality
Purpose
code_n.py
Python (Blender API)… See the full description on the dataset page: https://huggingface.co/datasets/MaxRondelli/BlenderRAG.bleeding-edge-gameplay-sampleThis dataset contains 1024 60 second video clips of Bleeding Edge gameplay (75GB). The data has already been processed into the following format: 300x180 videos sampled at 10 fps.
Dataset Structure
Data Files
testing_dataset_part1.zip & testing_dataset_part2.zip – Contains all 1024 60 second trajectories used for our evaluation.
4 examples from the dataset:
FB[…].npz – .npz file (described below)
FB[…].mp4 – 60 seconds .mp4 video of the images from the .npz file.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bleeding-edge-gameplay-sample.clipbench-blending
CLIPBench-Blending
Code, cached features and complete per-cell records for The Blending Ratio Is
Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot
Adaptation of Vision–Language Models (Liangzhi Li, Bowen Wang, Yiming Qian,
Thorsten Neumann, Xia Xie, and Guangshun Li; corresponding author Guangshun Li).
A large family of few-shot adaptation methods for vision–language models
classifies with a convex combination of the zero-shot text prototype and the
mean… See the full description on the dataset page: https://huggingface.co/datasets/Liangzhi-Li/clipbench-blending.blendedmvs_processeddolma-blend-gpt2MIC_sampledNemotron-RL-Super-Training-Blends
Dataset Description:
Nemotron-3-Super-RL-Training-Blends contains the dataset blends used to train the Nemotron-3-Super-120B-A12B model. RL training for the Nemotron-3-Super-120B-A12B model is done in 6 stages: RLVR 1, RLVR 2, RLVR 3, SWE 1, SWE 2, and RLHF. The blends for each stage consist of data from various datasets, which we detail below. The percentages in parentheses indicate the mixing ratios of the dataset components. Note that the model was also trained on additional data… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Super-Training-Blends.blenderbench-direct-results
BlenderBench Direct reproduction artifacts
This dataset preserves artifacts and provenance for an independent, community-run reproduction of the public BlenderBench task set. It is not an official Blender Foundation product, official BlenderBench submission, or leaderboard result.
Experiment
Dataset: DietCoke4671/BlenderBench revision 203e4d325e9438ca55b29bdfc4f6a90842d74e68
Attribution: DietCoke4671 and contributors, CC BY 4.0
Generation model: gpt-6-astra
Codex… See the full description on the dataset page: https://huggingface.co/datasets/michaelgold/blenderbench-direct-results.Hoyoverse_MapsNItems_BlenderblendBLEnD-Vis
BLEnD-Vis
BLEnD-Vis is a benchmark for evaluating vision-language models (VLMs) on culturally grounded multiple-choice questions, including a text-only setting and a visual setting with generated images.
Paper: https://arxiv.org/abs/2510.11178
Dataset repo: https://huggingface.co/datasets/Incomple/BLEnD-Vis
Code: https://github.com/Social-AI-Studio/BLEnD-Vis
Source
BLEnD-Vis is derived from the BLEnD dataset on Hugging Face (nayeon212/BLEnD).
What is in… See the full description on the dataset page: https://huggingface.co/datasets/Incomple/BLEnD-Vis.Nemotron-3-Nano-RL-Training-Blend
Dataset Description:
Nemotron-3-Nano-RL-Training-Blend is a curated dataset blend used to train the Nemotron-3-Nano-30B-A3B model. The blend consists of the following component datasets, with mixing ratios shown in parentheses:
nvidia/Nemotron-RL-instruction_following (0.17)
nvidia/Nemotron-RL-knowledge-mcqa (0.20)
nvidia/Nemotron-RL-agent-workplace_assistant (0.10)
nvidia/Nemotron-RL-instruction_following-structured_outputs (0.05)
nvidia/Nemotron-RL-coding-competitive_coding… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-3-Nano-RL-Training-Blend.Competitive-Programming-python-blend
Dataset Card for Competitive-Programming-python-blend
Summary
Competitive-Programming-python-blend is a mixed supervised fine-tuning dataset centered on competitive programming, code reasoning, and instruction-style problem solving. The blend is Python-first, but it also keeps a small amount of C++, agentless SWE, and reasoning-oriented chat supervision to broaden training coverage.
The current release is published as a single HF-friendly JSONL file, clean.jsonl.… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Competitive-Programming-python-blend.Nemotron-RL-Lightning-Training-Blend
Dataset Description:
This dataset provides the training-data blend used for the Reinforcement Learning with Verifiable Rewards (RLVR) stage of the public Nemotron-3.5-Lightning post-training recipe. The blend is consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. See the recipe for how the blend is used.
The blend mixes NVIDIA-released datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Lightning-Training-Blend.NOAA-PIFSC-ESD-CORAL-Bleaching-Dataset
Dataset Card for NOAA-ESD-CORAL-Bleaching Classification Dataset v1
Overview
For the development of machine learning models to classify coral health, specifically identifying healthy hard coral (CORAL) and bleached hard coral (CORAL_BL).This dataset contains underwater imagery collected by NOAA's Ecosystem Sciences Division (ESD) and other benthic surveys.
Labels
Label
Name
Functional Group
CORAL
Healthy Hard Coral
Hard Coral
CORAL_BL
Bleached… See the full description on the dataset page: https://huggingface.co/datasets/NMFS-OSI/NOAA-PIFSC-ESD-CORAL-Bleaching-Dataset.
