datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BLEnD
BLEnD
This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track).
24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.Nemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.Competitive-Programming-python-blend
Dataset Card for Competitive-Programming-python-blend
Summary
Competitive-Programming-python-blend is a mixed supervised fine-tuning dataset centered on competitive programming, code reasoning, and instruction-style problem solving. The blend is Python-first, but it also keeps a small amount of C++, agentless SWE, and reasoning-oriented chat supervision to broaden training coverage.
The current release is published as a single HF-friendly JSONL file, clean.jsonl.… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Competitive-Programming-python-blend.mix-instruct
MixInstruct
Introduction
This is the official realease of dataset MixInstruct for project LLM-Blender.
This dataset contains 11 responses from the current popular instruction following-LLMs that includes:
Stanford Alpaca
FastChat Vicuna
Dolly V2
StableLM
Open Assistant
Koala
Baize
Flan-T5
ChatGLM
MOSS
Moasic MPT
We evaluate each response with auto metrics including BLEU, ROUGE, BERTScore, BARTScore. And provide pairwise comparison results by prompting ChatGPT for the… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/mix-instruct.Mentor_Stage1BlendNet
📚 BlendNet
The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples.
For more details, please visit our GitHub repository or refer to our arXiv paper.
📖 Citation
@misc{du2024blenderllmtraininglargelanguage,
title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement},
author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/BlendNet.blender-3d-models
Blender 3D Models Database
This repository serves as a database of custom-made 3D models generated programmatically using Blender and uploaded to Hugging Face.
Models Included
⚔️ Stylized Fantasy Sword (stylized_sword.glb) - View Model Page
🛡️ Stylized Fantasy Shield (stylized_shield.glb) - View Model Page
🪄 Stylized Wizard's Staff (stylized_staff.glb) - View Model Page
📦 Stylized Treasure Chest (stylized_chest.glb) - View Model Page
🏹 Stylized Bow… See the full description on the dataset page: https://huggingface.co/datasets/abersbail/blender-3d-models.BlendX
BlendX : Complex Multi-Intent Detection with Blended Patterns
Official Repository for "BlendX : Complex Multi-Intent Detection with Blended Patterns." [Paper(ACL Anthology)] [Paper(arXiv)]
Yejin Yoon, Jungyeon Lee, Kangsan Kim, Chanhee Park and Taeuk Kim. Accepted to LREC-COLING2024 long paper.
Dataset Structure
./
├── v1.0/
│ ├── BlendX/
│ │ ├── BlendATIS/
│ │ ├── BlendBanking77/
│ │ ├── BlendCLINC150/
│ │ └── BlendSNIPS/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/HYU-NLP/BlendX.blender-dataset
blender-dataset
Dataset generated with DeepFabric.
kali_linux_toolkit_dataset
Kali Linux Tools Dataset
A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links.
This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers.
📁 Dataset Format
Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure… See the full description on the dataset page: https://huggingface.co/datasets/bleondubos/kali_linux_toolkit_dataset.perfectblend-smoltalk-chinese-large-blend-regenMentor_Stage2BlendNet
📚 BlendNet
The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples.
For more details, please visit our GitHub repository or refer to our arXiv paper.
📖 Citation
@misc{du2024blenderllmtraininglargelanguage,
title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement},
author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/spaudwal/BlendNet.blenderllm-v2-polyhaven-dataset
BlenderLLM v2 - Poly Haven Training Dataset
Fine-tuning dataset for BlenderLLM to use local Poly Haven library (2,194 assets).
Purpose
BlenderLLM v1 only generates primitive-based scripts. This dataset teaches it to:
Load models from local .blend files (426 models)
Apply HDRIs for realistic lighting (963 HDRIs)
Apply textures with PBR materials (805 textures)
Compose scenes combining real assets + primitives
Search/list available assets by category
Handle errors (missing… See the full description on the dataset page: https://huggingface.co/datasets/LiM-De/blenderllm-v2-polyhaven-dataset.typescript-chunks
typescript-chunks
A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol.
Processing
Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types.
FunctionDeclaration ---- 8205
ArrowFunction --------- 33890
ClassDeclaration ------- 5325
InterfaceDeclaration -- 12884
EnumDeclaration --------- 518
TypeAliasDeclaration --- 3580
MethodDeclaration ----- 24713
Leading comments are added to the… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-chunks.audio_to_blendshapes_maindeepfabric-blender-mcp
deepfabric-blender-mcp
Dataset generated with DeepFabric.
perfect-blend-gptoss-20B-1Mbless_luna_5kblender_duplicates
Dataset Card for Dataset Name
Contains reduced description of issues reported at https://projects.blender.org/blender/blender/issues and points to duplicate issues in order to categorize similarity.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Each report has been shortened by removing frequently repeated texts such as System Information, Blender Version… See the full description on the dataset page: https://huggingface.co/datasets/mano-wii/blender_duplicates.blended_skill_talk_ru
Dataset Card for "blended_skill_talk"
Dataset Summary
Russian version of the Blended Skill Talk dataset. Each replica was translated separately using a paid translator. A dataset of 7k conversations explicitly designed to exhibit multiple conversation modes: displaying personality, having empathy, and demonstrating knowledge.
Dataset Structure
Data Instances
An example of 'train' looks as follows.
{
"personas": ["мне все время звонит женщина."… See the full description on the dataset page: https://huggingface.co/datasets/artemsnegirev/blended_skill_talk_ru.kiswahili-ai-blended
Kiswahili AI — Swahili Instruction Model
AutoScientist Challenge 2026 — Language Category
Overview
Kiswahili AI is a Swahili instruction-tuned language model fine-tuned from Llama-4-Scout-17B-16E-Instruct (109B MoE) using AutoScientist by Adaption Labs. It combines 4 public Swahili datasets into a unified instruction dataset (~52K rows), processes them through Adaptive Data for quality enhancement, and trains via AutoScientist's closed-loop co-optimization.… See the full description on the dataset page: https://huggingface.co/datasets/NabajyotiPathak/kiswahili-ai-blended.audio_to_blendshapes_testBlendNet
📚 BlendNet
The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples.
For more details, please visit our GitHub repository or refer to our arXiv paper.
📖 Citation
@misc{du2024blenderllmtraininglargelanguage,
title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement},
author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/FlandreScarlet123/BlendNet.glossar_hackathon_blendCredits & Attribution: This dataset is a curated derivative blend utilizing slices from tasksource/social-chemistry-101 and the official NVIDIA Nemotron Post-Training dataset catalog (licensed under CC-BY-4.0 and ODC-By).
meg_converted_prompts_new_blended_split1qwen3-30b-0624-tooluse-blend-spare-games-envs
qwen3-30B-A3B-Instruct-0624-tooluse-blend — generated environments
Environments generated by the SPARE proposer during training run
09p118sw (qwen3-30B-A3B-Instruct-0624-tooluse-blend), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
243
Steps covered
9 (step 0–161)
With recovered skill
243
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0624-tooluse-blend-spare-games-envs.Novachrono-Reasoning-Blend-v1
🧠 Novachrono-Reasoning-Blend-v1
Novachrono-Reasoning-Blend-v1 is a large-scale, multi-source instruction dataset designed for training and evaluating reasoning-capable language models. The dataset contains structured instructions, intermediate reasoning annotations, and high-quality final responses across a diverse range of tasks and domains.
Built with a strong emphasis on clarity, consistency, and practical usefulness, this dataset is intended for instruction tuning, alignment… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/Novachrono-Reasoning-Blend-v1.perfectblend-smoltalk-chinese-blend-regencode-blend
