CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01databricks /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.textquestion-answering10K<n<100K1.2k likes62k downloads3y agoHugging Face02llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes43k downloads2y agoHugging Face03open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads15h agoHugging Face04RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face05johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M3 likes6.8k downloads8mo agoHugging Face06Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.7k downloads1y agoHugging Face07amitbcp /docinsights-2026-shared-task-data DocInsights 2026 Shared Task: DocSem Document-grounded quantitative reasoning with evidence attribution DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI. Workshop shared task | Source repository | Submission portal | Participant guide Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.documentquestion-answering1K<n<10K0 likes4.8k downloads21d agoHugging Face08yongxin2020 /TempPerturb-Eval-data TempPerturb-Eval-data Summary TempPerturb-Eval-data is the released output dataset for TempPerturb-Eval, a benchmark for analyzing the robustness of Retrieval-Augmented Generation (RAG) systems under both internal variation and external perturbation. This is an evaluation-artifact dataset: it stores model outputs and experiment metadata for controlled robustness analysis, rather than a new QA training corpus. The release covers: 5 models 11 temperatures from 0.0 to 2.0 4… See the full description on the dataset page: https://huggingface.co/datasets/yongxin2020/TempPerturb-Eval-data.textquestion-answering10K<n<100K1 likes3.6k downloads7mo agoHugging Face09lhpku20010120 /Data-Prep-Bench Data-Prep-Bench This repository contains the data presented in DataPrep-Bench: Benchmarking LLMs as Training Data Preparators. Code: https://github.com/OpenDCAI/Data-Preparation-Bench Dataset Overview This dataset is a comprehensive resource built for Supervised Fine-Tuning (SFT) and evaluation of Large Language Models (LLMs), covering six domains: Finance, Medicine, Law, Mathematics, Science, and General. A key feature of this dataset is that we employed 12… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Data-Prep-Bench.texttext-generation1M<n<10M1 likes3k downloads2mo agoHugging Face10Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes2.7k downloads2h agoHugging Face11daman1209arora /jeebench JEEBench(EMNLP 2023) Repository for the code and dataset for the paper: "Have LLMs Advanced Enough? A Harder Problem Solving Benchmark For Large Language Models" accepted in EMNLP 2023 as a Main conference paper. https://aclanthology.org/2023.emnlp-main.468/ Citation If you use our dataset in your research, please cite it using the following @inproceedings{arora-etal-2023-llms, title = "Have {LLM}s Advanced Enough? A Challenging Problem Solving Benchmark For Large… See the full description on the dataset page: https://huggingface.co/datasets/daman1209arora/jeebench.textquestion-answeringn<1K8 likes2.4k downloads3y agoHugging Face12minkyuchoi /Temporal-Logic-Video-Dataset Temporal Logic Video (TLV) Dataset Temporal Logic Video (TLV) Dataset Synthetic and real video dataset with temporal logic annotation Explore the GitHub » NSVS-TL Project Webpage · NSVS-TL Source Code Overview The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.tabularquestion-answeringn<1K1 likes2.4k downloads2y agoHugging Face13PolarSeeker /OpenSeeker-v1-Data OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data OpenSeeker is an open-source search agent system that democratizes access to frontier search capabilities by fully open-sourcing its training data. We fine-tuned Qwen3-30B-A3B-Thinking-2507 with 11.7K training examples and achieved state-of-the-art performance on frontier search benchmarks: Highlights Superior performance on search agent benchmarks: 48.4 on BrowseComp-ZH, 29.5 on… See the full description on the dataset page: https://huggingface.co/datasets/PolarSeeker/OpenSeeker-v1-Data.textquestion-answering10K<n<100K54 likes2.3k downloads6mo agoHugging Face14liarliar /Daily-OmniThis is the official dataset for Daily-Omni. Check code repository for instructions. textquestion-answering1K<n<10K5 likes2.3k downloads1y agoHugging Face15yongchao98 /R1-Code-Interpreter-Data R1-Code-Interpreter: Training LLMs to Reason with Code via Supervised and Reinforcement Learning Our code is based on Llama-factory/VeRL/Search-R1 for the SFT and RL training and SymBench/BIG-Bench-Hard/reasoning-gym for datasets/benchmarks of reasoning/planning tasks. 📝 Introduction R1-Code-Interpreter is the first framework to train LLMs for step-by-step code reasoning using multi-turn supervised fine-tuning and reinforcement learning. By curating 144 diverse… See the full description on the dataset page: https://huggingface.co/datasets/yongchao98/R1-Code-Interpreter-Data.textquestion-answering1K<n<10K2 likes2.2k downloads1y agoHugging Face16desearch /dataset Desearch Benchmark Questions Fresh, self-contained benchmark questions for evaluating web and X (Twitter) search. Regenerated daily from recent news and tweets. Each question is answerable from public sources within a dated window — there are no answer keys or source URLs in the public data, so systems have to actually search rather than recall. Subsets Path Lane Built from questions/ Web / news Recent news articles (RSS + news sitemaps) x/ X /… See the full description on the dataset page: https://huggingface.co/datasets/desearch/dataset.textquestion-answering100K<n<1M0 likes2k downloads7h agoHugging Face17snfacademy /personal-trainer-ausbildung-ki-datensatz SNFA Personal Trainer Ausbildung KI-Datensatz Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz. Inhalt Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.textquestion-answeringn<1K0 likes1.3k downloads2mo agoHugging Face18uriel /Maathis_Ohada_datasettextquestion-answering1K<n<10K2 likes1.3k downloads3y agoHugging Face19risenyard /egms-qa-dataset EGMS-QA Dataset Prepared EGMS displacement tiles, encoder tokens, task labels, reference tables, and natural-language QA records for 10,000 overlapping 7 km tiles. This card describes the available data, file formats, and download options. Data access Data needed Files to download Details Published QA records train.jsonl, validation.jsonl, test.jsonl QA loading example Encoder inputs Source tiles, metadata Encoder data Translator inputs Token cache… See the full description on the dataset page: https://huggingface.co/datasets/risenyard/egms-qa-dataset.textquestion-answering100K<n<1M1 likes1.1k downloads18d agoHugging Face20FreedomIntelligence /Medical-R1-Distill-Data Introduction This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1. The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese. The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.textquestion-answering10K<n<100K77 likes948 downloads2y agoHugging Face21Congliu /Chinese-DeepSeek-R1-Distill-data-110k 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face&nbsp;&nbsp; | &nbsp;&nbsp;🤖 ModelScope &nbsp;&nbsp; | &nbsp;&nbsp;🚀 Github &nbsp;&nbsp; | &nbsp;&nbsp;📑 Blog 注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。 该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.tabulartext-generation100K<n<1M789 likes817 downloads2y agoHugging Face22OpenGVLab /InternVL-Chat-V1-2-SFT-Data Data Card for InternVL-Chat-V1-2-SFT-Data Overview Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT. Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.imagevisual-question-answering100K<n<1M29 likes761 downloads2y agoHugging Face23talmahmud /my_dataset_repotextquestion-answering10K<n<100K0 likes702 downloads1y agoHugging Face24ChatTSRepo /ChatTS-Training-Dataset ChatTS-Training Data This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model. Datasets align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256. align_random: Alignment training dataset with random sequence lengths between 64 and 1024. sft: SFT dataset generated with Time Series Evol-Instruct. ift: Instruction following dataset. dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.textquestion-answering100K<n<1M14 likes652 downloads1y agoHugging Face25Snowflake /dare-bench DARE-Bench [ICLR 2026] DARE-Bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science Fan Shu1, Yite Wang2, Ruofan Wu1, Boyi Liu2, Zhewei Yao2, Yuxiong He2, Feng Yan1 1University of Houston   2Snowflake AI Research 🔎 Overview DARE-Bench (ICLR 2026) is a benchmark for evaluating LLM agents on data science tasks, focusing on modeling and instruction fidelity. This Hugging Face repository provides a selected subset of the full benchmark for public release.… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/dare-bench.texttext-generation1K<n<10K6 likes604 downloads7mo agoHugging Face26snfacademy /snfa-ernaehrungscoach-ki-dataset SNFA Ernährungscoach KI-Datensatz Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Ernährungscoaching, Online-Ausbildung, Beratung, Berufspraxis und Selbstständigkeit in der Schweiz. Inhalt Die Datei snfa_ernaehrungscoach_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei und Quelle. Die ursprünglichen… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/snfa-ernaehrungscoach-ki-dataset.textquestion-answeringn<1K0 likes542 downloads2mo agoHugging Face27utsavm /NSFW_Chat_Dataset 💕 Spicy AI GF Chat Dataset 🔥 🚨 18+ Only! NSFW & Spicy Content Ahead 🚨 Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you. 📜 What’s Inside? This dataset features two columns: input → Boyfriend’s dialogue (aka what YOU say 😉) output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.textquestion-answering1K<n<10K13 likes538 downloads2y agoHugging Face28reloading0101 /threat-intelligence-dataset Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.texttext-generation10K<n<100K10 likes523 downloads3mo agoHugging Face29Kylan12 /Synthetic-AI-ML-Dataset Synthetic-AI-ML-Dataset Synthetic Q&A dataset on AI and Machine Learning Dataset Details Metric Value Topic AI and Machine Learning Total Q&A Pairs 14021 Valid Pairs 14021 Provider/Model ollama/gpt-oss:120b Generation Cost Metric Value Prompt Tokens 14,941,957 Completion Tokens 17,159,263 Total Tokens 32,101,220 GPU Energy 12.9628 kWh Sources This dataset was generated from 474 scholarly papers: #… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.textquestion-answering10K<n<100K2 likes473 downloads6mo agoHugging Face30AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes437 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.