CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01uilab /BLEnD BLEnD This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track). 24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.textquestion-answering100K<n<1M16 likes1.6k downloads10d agoHugging Face02nvidia /Nemotron-RL-Ultra-Training-Blends Dataset Description: This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used. The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.tabulartext-generation10K<n<100K19 likes1.4k downloads2mo agoHugging Face03Jackrong /Competitive-Programming-python-blend Dataset Card for Competitive-Programming-python-blend Summary Competitive-Programming-python-blend is a mixed supervised fine-tuning dataset centered on competitive programming, code reasoning, and instruction-style problem solving. The blend is Python-first, but it also keeps a small amount of C++, agentless SWE, and reasoning-oriented chat supervision to broaden training coverage. The current release is published as a single HF-friendly JSONL file, clean.jsonl.… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Competitive-Programming-python-blend.texttext-generation10K<n<100K21 likes510 downloads7mo agoHugging Face04llm-blender /mix-instruct MixInstruct Introduction This is the official realease of dataset MixInstruct for project LLM-Blender. This dataset contains 11 responses from the current popular instruction following-LLMs that includes: Stanford Alpaca FastChat Vicuna Dolly V2 StableLM Open Assistant Koala Baize Flan-T5 ChatGLM MOSS Moasic MPT We evaluate each response with auto metrics including BLEU, ROUGE, BERTScore, BARTScore. And provide pairwise comparison results by prompting ChatGPT for the… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/mix-instruct.texttext-generation100K<n<1M38 likes356 downloads3y agoHugging Face05BleachNick /Mentor_Stage1tabular1M<n<10M0 likes348 downloads1y agoHugging Face06FreedomIntelligence /BlendNet 📚 BlendNet The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples. For more details, please visit our GitHub repository or refer to our arXiv paper. 📖 Citation @misc{du2024blenderllmtraininglargelanguage, title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement}, author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/BlendNet.tabular10K<n<100K12 likes223 downloads2y agoHugging Face07abersbail /blender-3d-models Blender 3D Models Database This repository serves as a database of custom-made 3D models generated programmatically using Blender and uploaded to Hugging Face. Models Included ⚔️ Stylized Fantasy Sword (stylized_sword.glb) - View Model Page 🛡️ Stylized Fantasy Shield (stylized_shield.glb) - View Model Page 🪄 Stylized Wizard's Staff (stylized_staff.glb) - View Model Page 📦 Stylized Treasure Chest (stylized_chest.glb) - View Model Page 🏹 Stylized Bow… See the full description on the dataset page: https://huggingface.co/datasets/abersbail/blender-3d-models.3dn<1K0 likes206 downloads4mo agoHugging Face08HYU-NLP /BlendX BlendX : Complex Multi-Intent Detection with Blended Patterns Official Repository for "BlendX : Complex Multi-Intent Detection with Blended Patterns." [Paper(ACL Anthology)] [Paper(arXiv)] Yejin Yoon, Jungyeon Lee, Kangsan Kim, Chanhee Park and Taeuk Kim. Accepted to LREC-COLING2024 long paper. Dataset Structure ./ ├── v1.0/ │ ├── BlendX/ │ │ ├── BlendATIS/ │ │ ├── BlendBanking77/ │ │ ├── BlendCLINC150/ │ │ └── BlendSNIPS/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/HYU-NLP/BlendX.texttext-classification100K<n<1M3 likes181 downloads1y agoHugging Face09nolabs /blender-dataset blender-dataset Dataset generated with DeepFabric. textn<1K0 likes86 downloads9mo agoHugging Face10bleondubos /kali_linux_toolkit_dataset Kali Linux Tools Dataset A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links. This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers. 📁 Dataset Format Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure… See the full description on the dataset page: https://huggingface.co/datasets/bleondubos/kali_linux_toolkit_dataset.textn<1K0 likes64 downloads9d agoHugging Face11zjxia /perfectblend-smoltalk-chinese-large-blend-regentext100K<n<1M0 likes62 downloads6mo agoHugging Face12BleachNick /Mentor_Stage2tabular1M<n<10M0 likes51 downloads1y agoHugging Face13spaudwal /BlendNet 📚 BlendNet The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples. For more details, please visit our GitHub repository or refer to our arXiv paper. 📖 Citation @misc{du2024blenderllmtraininglargelanguage, title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement}, author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/spaudwal/BlendNet.tabular10K<n<100K0 likes50 downloads26d agoHugging Face14LiM-De /blenderllm-v2-polyhaven-dataset BlenderLLM v2 - Poly Haven Training Dataset Fine-tuning dataset for BlenderLLM to use local Poly Haven library (2,194 assets). Purpose BlenderLLM v1 only generates primitive-based scripts. This dataset teaches it to: Load models from local .blend files (426 models) Apply HDRIs for realistic lighting (963 HDRIs) Apply textures with PBR materials (805 textures) Compose scenes combining real assets + primitives Search/list available assets by category Handle errors (missing… See the full description on the dataset page: https://huggingface.co/datasets/LiM-De/blenderllm-v2-polyhaven-dataset.textn<1K0 likes49 downloads6mo agoHugging Face15bleugreen /typescript-chunks typescript-chunks A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol. Processing Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. FunctionDeclaration ---- 8205 ArrowFunction --------- 33890 ClassDeclaration ------- 5325 InterfaceDeclaration -- 12884 EnumDeclaration --------- 518 TypeAliasDeclaration --- 3580 MethodDeclaration ----- 24713 Leading comments are added to the… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-chunks.texttext-classification10K<n<100K3 likes45 downloads3y agoHugging Face16sanjuhs /audio_to_blendshapes_maintextn<1K0 likes42 downloads1y agoHugging Face17nolabs /deepfabric-blender-mcp deepfabric-blender-mcp Dataset generated with DeepFabric. text1K<n<10K2 likes40 downloads9mo agoHugging Face18zhuyksir /perfect-blend-gptoss-20B-1Mtext100K<n<1M1 likes34 downloads1y agoHugging Face19h3inkr /bless_luna_5ktext10K<n<100K0 likes31 downloads19d agoHugging Face20mano-wii /blender_duplicates Dataset Card for Dataset Name Contains reduced description of issues reported at https://projects.blender.org/blender/blender/issues and points to duplicate issues in order to categorize similarity. This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Each report has been shortened by removing frequently repeated texts such as System Information, Blender Version… See the full description on the dataset page: https://huggingface.co/datasets/mano-wii/blender_duplicates.texttext-classification1K<n<10K0 likes27 downloads3y agoHugging Face21artemsnegirev /blended_skill_talk_ru Dataset Card for "blended_skill_talk" Dataset Summary Russian version of the Blended Skill Talk dataset. Each replica was translated separately using a paid translator. A dataset of 7k conversations explicitly designed to exhibit multiple conversation modes: displaying personality, having empathy, and demonstrating knowledge. Dataset Structure Data Instances An example of 'train' looks as follows. { "personas": ["мне все время звонит женщина."… See the full description on the dataset page: https://huggingface.co/datasets/artemsnegirev/blended_skill_talk_ru.text1K<n<10K1 likes26 downloads3y agoHugging Face22NabajyotiPathak /kiswahili-ai-blended Kiswahili AI — Swahili Instruction Model AutoScientist Challenge 2026 — Language Category Overview Kiswahili AI is a Swahili instruction-tuned language model fine-tuned from Llama-4-Scout-17B-16E-Instruct (109B MoE) using AutoScientist by Adaption Labs. It combines 4 public Swahili datasets into a unified instruction dataset (~52K rows), processes them through Adaptive Data for quality enhancement, and trains via AutoScientist's closed-loop co-optimization.… See the full description on the dataset page: https://huggingface.co/datasets/NabajyotiPathak/kiswahili-ai-blended.text10K<n<100K0 likes25 downloads3mo agoHugging Face23sanjuhs /audio_to_blendshapes_testtextn<1K0 likes22 downloads1y agoHugging Face24FlandreScarlet123 /BlendNet 📚 BlendNet The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples. For more details, please visit our GitHub repository or refer to our arXiv paper. 📖 Citation @misc{du2024blenderllmtraininglargelanguage, title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement}, author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/FlandreScarlet123/BlendNet.tabular10K<n<100K1 likes22 downloads8mo agoHugging Face25deleeharris /glossar_hackathon_blendCredits & Attribution: This dataset is a curated derivative blend utilizing slices from tasksource/social-chemistry-101 and the official NVIDIA Nemotron Post-Training dataset catalog (licensed under CC-BY-4.0 and ODC-By). text1K<n<10K0 likes18 downloads4mo agoHugging Face26justfordata /meg_converted_prompts_new_blended_split1text100K<n<1M0 likes17 downloads2y agoHugging Face27msr-spare-1 /qwen3-30b-0624-tooluse-blend-spare-games-envs qwen3-30B-A3B-Instruct-0624-tooluse-blend — generated environments Environments generated by the SPARE proposer during training run 09p118sw (qwen3-30B-A3B-Instruct-0624-tooluse-blend), recovered from the spare-viz durable cache. The run's scratch directory no longer exists; this dataset is the surviving copy. Games 243 Steps covered 9 (step 0–161) With recovered skill 243 With hint 0 Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0624-tooluse-blend-spare-games-envs.textn<1K0 likes17 downloads1mo agoHugging Face28NovachronoAI /Novachrono-Reasoning-Blend-v1 🧠 Novachrono-Reasoning-Blend-v1 Novachrono-Reasoning-Blend-v1 is a large-scale, multi-source instruction dataset designed for training and evaluating reasoning-capable language models. The dataset contains structured instructions, intermediate reasoning annotations, and high-quality final responses across a diverse range of tasks and domains. Built with a strong emphasis on clarity, consistency, and practical usefulness, this dataset is intended for instruction tuning, alignment… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/Novachrono-Reasoning-Blend-v1.texttext-generation10K<n<100K2 likes15 downloads9mo agoHugging Face29zjxia /perfectblend-smoltalk-chinese-blend-regentext100K<n<1M0 likes14 downloads6mo agoHugging Face30PishangShedappp /code-blendtext1M<n<10M0 likes14 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.