CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TokenBender /code_instructions_122k_alpaca_styletext100K<n<1M80 likes5.6k downloads3y agoHugging Face02TheTokenFactory /sec-contracts-financial-extraction-instructions S&P 500 SEC Financial Extraction Instructions Dataset Summary 7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies: Split Examples Filing Type Description train 3,430 Exhibit 10 + DEF 14A Positive examples with validated outputs corrective 4,253 Exhibit 10 + DEF 14A Corrective, rescued, and negative examples Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.texttext-generation10K<n<100K1 likes3.2k downloads6mo agoHugging Face03mesolitica /Malay-Dialect-Instructions Malay dialect instruction including coding Negeri Sembilan QA public transport QA, Coding CUDA coding, Kedah QA infra QA, Coding Rust coding, Kelantan QA Najib Razak QA, Coding Go coding, Perak QA Anwar Ibrahim QA, Coding SQL coding, Pahang QA Pendatang asing QA, Coding Typescript coding, Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.texttext-generation10K<n<100K6 likes2.5k downloads2y agoHugging Face04AdaptLLM /food-visual-instructions Adapting Multimodal Large Language Models to Domains via Post-Training (EMNLP 2025) This repos contains the food visual instructions for post-training MLLMs in our paper: On Domain-Specific Post-Training for Multimodal Large Language Models. The main project page is: Adapt-MLLM-to-Domains Data Information Using our visual instruction synthesizer, we generate visual instruction tasks based on the image-caption pairs from extended Recipe1M+ dataset. These synthetic… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/food-visual-instructions.imagevisual-question-answering100K<n<1M3 likes1.2k downloads1y agoHugging Face05jayelm /natural-instructionsPreprocessed version of Super-Natural-Instructions from https://github.com/allenai/natural-instructions/tree/master/splits. The same inputs may appear with different outputs, thus to avoid duplicate inputs, you can deduplicate by the id or the inputs field. This is modified from https://huggingface.co/datasets/Muennighoff/natural-instructions with a few improvements: Adds positive/negative examples, outputs, explanations for each task, to support different task definitions. Adds an "eval"… See the full description on the dataset page: https://huggingface.co/datasets/jayelm/natural-instructions.textother1M<n<10M4 likes1k downloads4y agoHugging Face06mesolitica /instructions-pair-miningtext100K<n<1M2 likes975 downloads3y agoHugging Face07pegah-a /small-natural-instructionstext100K<n<1M1 likes899 downloads3y agoHugging Face08jhu-clsp /core17-instructionstexttext-retrieval10K<n<100K2 likes593 downloads7mo agoHugging Face09malaysia-ai /mosaic-instructions Mosaic format for instructions dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-instructions.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-instructions load it, from streaming import LocalDataset… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-instructions.textn<1K1 likes561 downloads3y agoHugging Face10jhu-clsp /robust04-instructionstexttext-retrieval100K<n<1M2 likes538 downloads7mo agoHugging Face11universalner /uner_llm_instructions Dataset Card for Universal NER v1 in the Aya format This dataset is a format conversion from its original v1 format into the Aya instruction format and it's released here under the same CC-BY-SA 4.0 license and conditions. It contains data in multiple languages and this version is intended for multi-lingual LLM construction/tuning. The dataset contains different subsets and their dev/test/train splits, depending on language. Citation If you utilize this dataset version… See the full description on the dataset page: https://huggingface.co/datasets/universalner/uner_llm_instructions.texttoken-classification10K<n<100K2 likes527 downloads3y agoHugging Face12jhu-clsp /news21-instructionstexttext-retrieval10K<n<100K1 likes522 downloads7mo agoHugging Face13cfahlgren1 /react-code-instructions React Code Instructions Popular Queries Number of instructions by Model Unnested Messages Instructions Added Per Day Dataset of Claude Artifact esque React Apps generated by Llama 3.1 70B, Llama 3.1 405B, and Deepseek Chat V3. Examples Virtual Fitness Trainer Website LinkedIn Clone iPhone Calculator Chipotle Waitlist Apple Store text10K<n<100K158 likes505 downloads2y agoHugging Face14axiong /pmc_llama_instructionsThis repo provides part of the dataset used for PMC-LLaMA-13B's instruction tuning. Data Size Link ChatDoctor 100K https://www.yunxiangli.top/ChatDoctor/ MedQA 10.2K https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options MedMCQA 183K https://huggingface.co/datasets/medmcqa PubmedQA 211K https://huggingface.co/datasets/pubmed_qa LiveQA 635 https://huggingface.co/datasets/truehealth/liveqa MedicationQA 690 https://huggingface.co/datasets/truehealth/medicationqa UMLS… See the full description on the dataset page: https://huggingface.co/datasets/axiong/pmc_llama_instructions.textquestion-answering100K<n<1M33 likes445 downloads3y agoHugging Face15Den4ikAI /russian_instructions_2June 10: Почищены криво переведенные примеры кода Добавлено >50000 человеческих примеров QA и инструкций Обновленная версия русского датасета инструкций и QA. Улучшения: 1. Увеличен размер с 40 мегабайт до 130 (60к сэмплов - 200к) 2. Улучшено качество перевода. Структура датасета: { "sample":[ "Как я могу улучшить свою связь между телом и разумом?", "Начните с разработки регулярной практики осознанности. 2. Обязательно практикуйте баланс на нескольких уровнях: физическом… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_instructions_2.text100K<n<1M27 likes426 downloads3y agoHugging Face16aarajbhattarai /law-instructions-dataset Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.texttext-generation1K<n<10K0 likes354 downloads11d agoHugging Face17heegyu /open-korean-instructions4가지 한국어 챗봇 학습용 데이터셋을 합쳐놓았습니다. 이중 ShareGPT 데이터는 멀티턴으로 되어있습니다. 데이터 생성 및 합치는 코드는 https://github.com/HeegyuKim/open-korean-instructions 여기를 참고하세요 이름 # 타입 KoAlpaca v1.0 52K 싱글턴 KoAlpaca v1.1 21K 싱글턴 ShareGPT DeepL 번역 620K(싱글턴), 84K(멀티턴) 멀티턴, 싱글턴 OIG-small-chip2-ko 210K 싱글턴 Korquad-Chat 9.6K 멀티턴, 지식기반 모든 데이터는 포멧이 통일되어 있습니다. <sys>, <usr>, <bot> 세가지 토큰과 줄넘김으로 화자를 구분합니다. korquad-chat 데이터의 경우, 유저와 봇이 서로를 호칭할 때는 <|bot|>, <|user|>로 되어있습니다. {"source": "koalpaca-v1.0", "text":… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/open-korean-instructions.text100K<n<1M25 likes320 downloads3y agoHugging Face18NickIBrody /python-code-instructions-85k Python Code Instructions - 85K Instruction-tuning dataset of Python functions paired with short natural-language instructions derived from repository docstrings. What changed in this release This release keeps the original public rows and format, but makes the dataset easier to use responsibly: exact duplicate rows were removed again using normalized instruction + output hashing deterministic train, validation, and test splits were added the dataset card now documents… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/python-code-instructions-85k.texttext-generation10K<n<100K1 likes318 downloads5mo agoHugging Face19Den4ikAI /russian_instructionsНовая версия: https://huggingface.co/datasets/Den4ikAI/russian_instructions_2 Русский датасет инструкций и QA. Структура датасета: { "dialogue":[ "Как я могу улучшить свою связь между телом и разумом?", "Начните с разработки регулярной практики осознанности. 2. Обязательно практикуйте баланс на нескольких уровнях: физическом, эмоциональном, умственном и духовном. 3. Свяжитесь с природой, когда это возможно - идите на прогулки или бегайте на улице, или просто сидите в парке и… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_instructions.text10K<n<100K21 likes281 downloads4y agoHugging Face20aarajbhattarai /rejected-agriculture-instructions-dataset Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-agriculture-instructions-dataset.texttext-generation10K<n<100K0 likes276 downloads14d agoHugging Face21aarajbhattarai /agriculture-instructions-dataset Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/agriculture-instructions-dataset.texttext-generation10K<n<100K0 likes269 downloads14d agoHugging Face22aarajbhattarai /unjudged-agriculture-instructions-dataset Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-agriculture-instructions-dataset.texttext-generationn<1K0 likes268 downloads14d agoHugging Face23dikshyamohanty /natural-instructions-sampletext10K<n<100K0 likes267 downloads3y agoHugging Face24paperbd /paper_instructions_300K-v1Loading will work as follows: Existing behavior # Loads the SFT dataset containing instruction, prompt, output load_dataset("paperbd/paper_instructions_300K-v1") Reasoning variant # Loads reasoning subset containing instruction, prompt, reasoning, output load_dataset( "paperbd/paper_instructions_300K-v1", "reasoning", split="train", ) Dataset Summary This dataset contains synthetic supervised fine-tuning data generated from academic… See the full description on the dataset page: https://huggingface.co/datasets/paperbd/paper_instructions_300K-v1.textquestion-answering100K<n<1M14 likes203 downloads4mo agoHugging Face25DonV1to /lego-instructions-public LEGO instruction PDF index This public dataset contains one curated US-Letter or language-neutral visual instruction PDF per current LEGO booklet. Explicit translated extras, obsolete asset revisions, corrupt files, and duplicate file contents are excluded. instruction-manifest.jsonl is the authoritative index. Each row records the set number, year, set name, booklet position, retained LEGO asset identifier, source URL, SHA-256 digest, byte size, and dataset-relative PDF path.… See the full description on the dataset page: https://huggingface.co/datasets/DonV1to/lego-instructions-public.documentvisual-question-answering1K<n<10K0 likes171 downloads1mo agoHugging Face26aarajbhattarai /unjudged-law-instructions-dataset Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-law-instructions-dataset.texttext-generationn<1K1 likes169 downloads11d agoHugging Face27Jerome-Young /OrthoTryOn-Instructions OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation Model Introduction We introduce OrthoTryOn, a unified and parameter-efficient framework for fashion image generation, designed to mitigate inter-task interference in shared adaptation and enable high-quality virtual try-on, garment reconstruction, and pose transfer within a single model. Its plug-and-play design can further extend to broader multi-task scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/Jerome-Young/OrthoTryOn-Instructions.text10K<n<100K1 likes164 downloads3mo agoHugging Face28vishnuOI /unity-dev-instructions Unity Developer Instructions A comprehensive instruction-tuning dataset for Unity game development, covering C# scripting, XR/VR development, physics, animation, rendering, UI Toolkit, and performance optimization. Dataset Summary Split Count Train 46,483 Test 2,446 Total 48,929 Data Sources | unity_docs | 40,496 | | stackoverflow | 6,071 | | github | 2,362 | Source breakdown: Source Count unity_docs 40,496 stackoverflow 6,071… See the full description on the dataset page: https://huggingface.co/datasets/vishnuOI/unity-dev-instructions.texttext-generation10K<n<100K10 likes158 downloads6mo agoHugging Face29sozercan /k8s-instructionsThis is a fork from https://huggingface.co/datasets/substratusai/k8s-instructions textn<1K5 likes155 downloads3y agoHugging Face30Pinkstack /LuauDev-instructions-SFT-preview LuauDev-SFT-PREVIEW THIS IS A PREVIEW VARIANT OF LUAUDEV. non preview: Pinkstack/LuauDev-instructions-SFT-full This is an SFT dataset meant for training Luau(Roblox's coding language) oriented large language models. Once the full version would be out it would be the biggest Luau instruction-style dataset ever released. These are the models which were used for data generation: (no specific order) DiffusionGemma 26B A4B Deepseek v4 Flash 0731 Nemotron 3 Ultra 550B A55B dots3 note… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstack/LuauDev-instructions-SFT-preview.text1K<n<10K2 likes153 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.