datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blended_skill_talk
Dataset Card for "blended_skill_talk"
Dataset Summary
A dataset of 7k conversations explicitly designed to exhibit multiple conversation modes: displaying personality, having empathy, and demonstrating knowledge.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 38.11 MB
Size of the generated dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ParlAI/blended_skill_talk.Unified-FeedbackCollections of pairwise feedback datasets.
openai/summarize_from_feedback
openai/webgpt_comparisons
Dahoas/instruct-synthetic-prompt-responses
Anthropic/hh-rlhf
lmsys/chatbot_arena_conversations
openbmb/UltraFeedback
argilla/ultrafeedback-binarized-preferences-cleaned
berkeley-nest/Nectar
Codes to reproduce the dataset: jdf-prog/UnifiedFeedback
Dataset formats
{
"id": "...",
"conv_A": [
{
"role": "user",
"content": "...",
},
{
"role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/Unified-Feedback.UltraEdit_500k
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/BleachNick/UltraEdit_500k.cua-blenderBLEnD
BLEnD
This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track).
24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.Nemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.BlenderBench
BlenderBench Dataset
Dataset Description
BlenderBench is a comprehensive benchmark dataset for evaluating models on 3D scene editing tasks in Blender. The dataset challenges agents to understand visual differences between initial and target scenes, then generate appropriate Blender Python code to transform the initial scene to match the target.
Key Features
27 instances across 3 difficulty levels
Multi-modal: Combines visual… See the full description on the dataset page: https://huggingface.co/datasets/DietCoke4671/BlenderBench.UltraEdit_Region_Based_100k
Bibtex citation
@misc{zhao2024ultraeditinstructionbasedfinegrainedimage,
title={UltraEdit: Instruction-based Fine-Grained Image Editing at Scale},
author={Haozhe Zhao and Xiaojian Ma and Liang Chen and Shuzheng Si and Rujie Wu and Kaikai An and Peiyu Yu and Minjia Zhang and Qing Li and Baobao Chang},
year={2024},
eprint={2407.05282},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2407.05282},
}
BlenderRAG
BlenderRAG Dataset
A dataset for 3D scene and object generation research. Each sample pairs a Blender Python script that procedurally generates a 3D object with a rendered preview image and a natural-language description.
Dataset Summary
The dataset is organized into two top-level scenes — indoor and outdoor — each containing a collection of objects. Every object is represented by three aligned modalities:
File
Modality
Purpose
code_n.py
Python (Blender API)… See the full description on the dataset page: https://huggingface.co/datasets/MaxRondelli/BlenderRAG.MIC_sampledblendBLEnD-Vis
BLEnD-Vis
BLEnD-Vis is a benchmark for evaluating vision-language models (VLMs) on culturally grounded multiple-choice questions, including a text-only setting and a visual setting with generated images.
Paper: https://arxiv.org/abs/2510.11178
Dataset repo: https://huggingface.co/datasets/Incomple/BLEnD-Vis
Code: https://github.com/Social-AI-Studio/BLEnD-Vis
Source
BLEnD-Vis is derived from the BLEnD dataset on Hugging Face (nayeon212/BLEnD).
What is in… See the full description on the dataset page: https://huggingface.co/datasets/Incomple/BLEnD-Vis.Competitive-Programming-python-blend
Dataset Card for Competitive-Programming-python-blend
Summary
Competitive-Programming-python-blend is a mixed supervised fine-tuning dataset centered on competitive programming, code reasoning, and instruction-style problem solving. The blend is Python-first, but it also keeps a small amount of C++, agentless SWE, and reasoning-oriented chat supervision to broaden training coverage.
The current release is published as a single HF-friendly JSONL file, clean.jsonl.… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Competitive-Programming-python-blend.mix-instruct
MixInstruct
Introduction
This is the official realease of dataset MixInstruct for project LLM-Blender.
This dataset contains 11 responses from the current popular instruction following-LLMs that includes:
Stanford Alpaca
FastChat Vicuna
Dolly V2
StableLM
Open Assistant
Koala
Baize
Flan-T5
ChatGLM
MOSS
Moasic MPT
We evaluate each response with auto metrics including BLEU, ROUGE, BERTScore, BARTScore. And provide pairwise comparison results by prompting ChatGPT for the… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/mix-instruct.Mentor_Stage1blends
Bangumi Image Base of Blend S
This is the image base of bangumi Blend S, we detected 16 characters, 1863 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters' preview:… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/blends.blendedmvsBlendNet
📚 BlendNet
The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples.
For more details, please visit our GitHub repository or refer to our arXiv paper.
📖 Citation
@misc{du2024blenderllmtraininglargelanguage,
title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement},
author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/BlendNet.blender-3d-models
Blender 3D Models Database
This repository serves as a database of custom-made 3D models generated programmatically using Blender and uploaded to Hugging Face.
Models Included
⚔️ Stylized Fantasy Sword (stylized_sword.glb) - View Model Page
🛡️ Stylized Fantasy Shield (stylized_shield.glb) - View Model Page
🪄 Stylized Wizard's Staff (stylized_staff.glb) - View Model Page
📦 Stylized Treasure Chest (stylized_chest.glb) - View Model Page
🏹 Stylized Bow… See the full description on the dataset page: https://huggingface.co/datasets/abersbail/blender-3d-models.task1418_bless_semantic_relation_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1418_bless_semantic_relation_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1418_bless_semantic_relation_classification.BlendX
BlendX : Complex Multi-Intent Detection with Blended Patterns
Official Repository for "BlendX : Complex Multi-Intent Detection with Blended Patterns." [Paper(ACL Anthology)] [Paper(arXiv)]
Yejin Yoon, Jungyeon Lee, Kangsan Kim, Chanhee Park and Taeuk Kim. Accepted to LREC-COLING2024 long paper.
Dataset Structure
./
├── v1.0/
│ ├── BlendX/
│ │ ├── BlendATIS/
│ │ ├── BlendBanking77/
│ │ ├── BlendCLINC150/
│ │ └── BlendSNIPS/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/HYU-NLP/BlendX.tailor_datasetIndspeech-Augmented-Dataset-10000-20000BlenderCAD2bleedingheart-pretrain-10MBleedingheart Pretrain Dataset
A collaboration between Kaleido and Newstar
We collected all the datasets we could find that are in Tagalog or any other Philippine dialect and put them in this repository.
This data will be used to train the Bleedingheart model.
Bleeding Heart is a stunning bird native to the island of Luzon in the Philippines. It is a medium-sized ground dove with a distinctive red patch of feathers on its chest, which gives it its name. The male's red patch is… See the full description on the dataset page: https://huggingface.co/datasets/NewstaR/bleedingheart-pretrain-10M.blender-dataset
blender-dataset
Dataset generated with DeepFabric.
BLEUBERI-Tulu3-50k[Paper] [HF Collection] [Code]
Authors: Yapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh, Chuan Li, Chris Tanner, Mohit Iyyer
Contact: yapeic@umd.edu
TLDR > We extend RLVR beyond easily verifiable domains like math and code to the more open-ended setting of general instruction following. Surprisingly, we find that BLEU—a simple n-gram matching metric—when paired with high-quality references from strong LLMs, achieves human agreement comparable to 8B and 27B reward models on Chatbot… See the full description on the dataset page: https://huggingface.co/datasets/yapeichang/BLEUBERI-Tulu3-50k.task1582_bless_hypernym_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1582_bless_hypernym_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1582_bless_hypernym_generation.Nemotron-3-Nano-RL-Training-Blend-prompt-only
Nemotron-3-Nano-RL-Training-Blend-prompt-only
Prompt-only extraction from nvidia/Nemotron-3-Nano-RL-Training-Blend.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-3-Nano-RL-Training-Blend-prompt-only.perfectblend-smoltalk-chinese-large-blend-regen
