CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K18 likes4.5k downloads10mo agoHugging Face02ASTRAI-labs /Pluto-Nano-1.0-Pretrain-v2 ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.tabulartext-generation10M<n<100M2 likes820 downloads3mo agoHugging Face03LiquidAI /ifstruct-v1.0 IFStruct v1.0 [!Note] 📝 Blog post: https://www.liquid.ai/blog/ifstruct-v1.0 💻 GitHub: https://github.com/Liquid4All/ifstruct IFStruct is a benchmark for structured-output compliance: can a model produce valid JSON/YAML that follows a requested schema, when the requirements are phrased the many different ways real users phrase them? It is scored without constrained decoding, and only the structure is judged (not content quality, extraction accuracy, or reasoning) so the… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/ifstruct-v1.0.texttext-generation1K<n<10K79 likes581 downloads3mo agoHugging Face04AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes426 downloads11mo agoHugging Face05Solstice-AI /Solace-1.0-Omnigated Project Solace The largest verified frontier-model distillation corpus ever released. 60 datasets · 7 frontier model families · 12,586,893 unique conversations · One file · Zero filler The short version This is synthetic data. The best kind of synthetic data. Every example was generated by a verified 2026 frontier model — GLM-5.2, Claude Fable 5, Mythos 5, GPT-5.6 Sol, GPT-5.5 Codex, DeepSeek V4 Pro 0813, Qwen 3.8-Max, and Kimi K3 — then… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-1.0-Omni.texttext-generation10M<n<100M6 likes423 downloads21d agoHugging Face06AI45Research /AgentDoG1.0-Training-Data AgentDoG1.0 Training Data [💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection] AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis. Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.texttext-generation1K<n<10K0 likes422 downloads4mo agoHugging Face07latam-gpt /LatamGPT-Corpus-1.0gated LatamGPT-Corpus-1.0 🌐 Language versions: English | Español | Português 🔗 Project links: Official LatamGPT website | Corpus dashboard 🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0. Dataset description Summary LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.imagetext-generation100M<n<1B8 likes378 downloads10d agoHugging Face08llm-jp /magpie-sft-v1.0 magpie-sft-v1.0 This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This is a dataset of instruction and response pairs created using the Magpie method. cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.texttext-generation100K<n<1M19 likes350 downloads2y agoHugging Face09khtsly /Luau-Coder-1.0-Preview-SFT Luau Coder 1.0 Preview SFT 🦭 This dataset is exceptionally high-quality supervised fine-tuning conversations for a highly capable coding model in Roblox Luau domain. It prioritize technical correctness, useful engineering judgment, realistic interaction, and efficient explanations over output volume. This dataset includes & covering: Multi-turns (4-10 turns) Dynamic CoT (length) Dynamic Interleaved Reasoning Long Context Session Q/A Review Debugging Bug Fix… See the full description on the dataset page: https://huggingface.co/datasets/khtsly/Luau-Coder-1.0-Preview-SFT.texttext-generation10K<n<100K1 likes309 downloads9d agoHugging Face10nhagar /dclm-baseline-1.0-parquet_urls Dataset Card for dclm-baseline-1.0-parquet_urls This dataset provides the URLs and top-level domains associated with training records in mlfoundations/dclm-baseline-1.0-parquet. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dclm-baseline-1.0-parquet_urls.texttext-generation1B<n<10B0 likes265 downloads1y agoHugging Face11taejoon89 /Ko-Agent-Trajectories-1.0 Ko-Agent-Trajectories-1.0 Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository under pipeline/, together with the API catalogue, the scenario templates and the complete prompt set. The card reports the completed human review study and the v1.1 artefacts (behaviour DPO config, per-item validation scores, manifest, filter asset). Korean edition: README.ko.md. TL;DR A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.tabulartext-generation100K<n<1M0 likes234 downloads2d agoHugging Face12OmniAICreator /Qiita-1.07MThis dataset contains 1,074,174 articles published on Qiita. tabulartext-classification1M<n<10M2 likes232 downloads1y agoHugging Face13wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to its EPR Labs source dataset. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.tabulartext-generation1M<n<10M0 likes229 downloads18d agoHugging Face14yuqing1207 /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K0 likes210 downloads9mo agoHugging Face15ZennyKenny /tactical-military-reasoning-v.1.0 Tactical Military Reasoning Dataset v1.0 A curated collection of 150 rich tactical military scenarios with LLM-generated reasoning strategies for both attacking and defending forces. 📝 Preface Oncologists do not study cancer because they love cancer and wish for it to occur more frequently. They study cancer to better understand its causes, progression, and consequences in order to therefore eradicate it from the earth more effectively. A distaste for something… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/tactical-military-reasoning-v.1.0.texttext-generationn<1K24 likes197 downloads1y agoHugging Face16ronaldocloud /cyberusecase-v1.0 Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level cybersecurity reasoning across vulnerability management, SOC alert triage, detection engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps. It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.texttext-generation10K<n<100K0 likes146 downloads3mo agoHugging Face17YouAIData /stem-reasoning-v1.0.0-ccbysa-001 YouAI Data — stem-reasoning-v1.0.0-ccbysa-001 Dataset Description YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 394 step-by-step reasoning chains and 569 instruction/response pairs across 332 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 364 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-001.texttext-generation1K<n<10K4 likes144 downloads5mo agoHugging Face18d0rj /ROMB-1.0 ♦ ROMB Русское описание и инструкция ROMB (Russian Olympiad Math Benchmark) evaluates models on Russian-language school olympiad mathematics. The test set contains 2552 text-only tasks: 1716 arithmetic/other tasks, 644 logic tasks, and 192 geometry tasks. Tasks have typed answers, answer-format notes, and per-task checking rules. The evaluator also supports configurable v3 runs: native thinking, optional JSON Schema constrained decoding, plain or \boxed{…} answers, and… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/ROMB-1.0.texttext-generation1K<n<10K1 likes142 downloads15d agoHugging Face19nanskong /ManipuriGPT-Corpus-v1.0 ManipuriGPT Corpus v1.0 ManipuriGPT Corpus v1.0 is a research-grade, multi-script, deduplicated, and quality-scored corpus specifically engineered for pretraining Manipuri (Meiteilon) language foundation models. Quick Summary Total Sequences: 147,956 Total Tokens (ManipuriGPT-Tokenizer-v1.0): 4,347,075 Total Characters: 16,019,401 Pipeline Version: 5.6 Release Version: v1.0.0 Build Timestamp: 2026-07-25T09:17:21.960438Z Primary Writing Systems… See the full description on the dataset page: https://huggingface.co/datasets/nanskong/ManipuriGPT-Corpus-v1.0.tabulartext-generation100K<n<1M0 likes139 downloads2mo agoHugging Face20khaimaitien /qa-expert-multi-hop-qa-V1.0 Dataset Card for QA-Expert-multi-hop-qa-V1.0 This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering. In total, this dataset contains 25.5k for training and 3.19k for evaluation. You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0 The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.textquestion-answering10K<n<100K8 likes120 downloads3y agoHugging Face21jumplander /JL-ActionBoundary-1K-v1.0.0 JL-ActionBoundary-1K v1.0.0 Counterfactual Ask–Inspect–Act–Defer supervision for coding agents JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code: ACT: the task is sufficiently specified for bounded repository work; INSPECT: missing information can be recovered from the repository; ASK: a material product decision belongs to the user; DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.tabulartext-classification1K<n<10K5 likes120 downloads2mo agoHugging Face22zait-ai /OpenOcypus-1.0 🪲 OpenOcypus-1.0 📖 Description OpenOcypus-1.0 is the first release in the OpenOcypus series — a collection of high‑quality SFT datasets designed to evolve over time. Future versions will introduce additional sources, refined filtering, and expanded task coverage. Total size: 1,157,428 examples. The dataset is designed to create a versatile assistant capable of: 🗣️ Engaging in natural conversations 🧮 Solving math problems with step‑by‑step explanations 💻… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/OpenOcypus-1.0.texttext-generation1M<n<10M0 likes108 downloads24d agoHugging Face23llm-bg /Tucan-BG-v1.0 Tucan-BG Dataset v1.0 Bilingual Function Calling Training Dataset for Bulgarian Language Models 🇧🇬 📄 Supporting the Tucan model series Paper: https://arxiv.org/abs/2506.23394 Overview 🚀 Tucan-BG-v1.0 is a bilingual (Bulgarian/English) dataset containing 10,035 conversations specifically designed for training language models in function calling and tool use capabilities. This dataset enables the development of AI agents capable of determining when to use… See the full description on the dataset page: https://huggingface.co/datasets/llm-bg/Tucan-BG-v1.0.texttext-generation10K<n<100K1 likes102 downloads1y agoHugging Face24pythainlp /han-instruct-dataset-v1.0 Dataset Card for "han-instruct-dataset-v1.0" The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset. Dataset Summary 🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. It collect the instruction following in Thai from many source. Many question are collect from Reference desk at Thai wikipedia. Data sources: Reference desk at Thai wikipedia. Law from justicechannel.org… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v1.0.texttext-generation1K<n<10K4 likes79 downloads5d agoHugging Face25AtomixLabs /OpenTopics-1.0-20K OpenTopics-1.0-20K What is this dataset? OpenTopics-1.0-20K is a collection of 20,003 topic names spanning a wide variety of subjects, including physics, medicine, history, law, engineering, and the arts. AtomixLabs built this dataset to help developers, researchers, and AI builders who need a large, organized list of topics. It works great for creating synthetic prompts, testing search systems, and training models to classify text. What is inside… See the full description on the dataset page: https://huggingface.co/datasets/AtomixLabs/OpenTopics-1.0-20K.tabulartext-classification10K<n<100K3 likes79 downloads2mo agoHugging Face26wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512 midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512 Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to Maestro, GiantMIDI, ATEPP, or MusicNet. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512.tabulartext-generation10K<n<100K0 likes75 downloads1mo agoHugging Face27AtomixLabs /OmniRoute-SFT-1.0-32K OmniRoute-SFT-1.0-32K AS OF 7/31/2026 The first large-scale supervised dataset for training LLM routing models. Every day, production AI systems face a deceptively simple question: which model should handle this query? A one-line cooking question doesn't need a 400-billion-parameter reasoning engine. A graduate-level proof doesn't belong on a chat-optimized 8B model. Routing gets this right — and saves orders of magnitude in compute — but until now, there has been no… See the full description on the dataset page: https://huggingface.co/datasets/AtomixLabs/OmniRoute-SFT-1.0-32K.texttext-classification10K<n<100K2 likes74 downloads28d agoHugging Face28pythainlp /thai_food_v1.0 Thai Food Recipe dataset v1.0 The Thai Food Recipe dataset is a collection of Thai recipes from old Thai books and social networks. List Book ตำรับอาหาร - เตื้อง สนิทวงศ์, ม.ร.ว., 2426-2510 - Work in process (ยังไม่ครบ) in v2.0 ผัดกะเพรา - “ทีมครัวเนื้อหอม” จ.ลำปาง สูตร "เกี๊ยวกุ้ง" License: cc0-1.0 texttext-generationn<1K9 likes72 downloads3y agoHugging Face29Kenotic-Labs /ATANTV1.0-corpus ATANT Narrative Test Corpus Automated Test for Acceptance of Narrative Truth, v1.0 The first open evaluation corpus for measuring continuity in AI systems: the ability to persist, update, disambiguate, and reconstruct meaningful context across time. Paper: ATANT: An Evaluation Framework for AI Continuity (arXiv:2604.06710) Standard repository: github.com/Kenotic-Labs/ATANT Author: Samuel Sameer Tanguturi Affiliation: Kenotic Labs Published: April 2026 Why this corpus… See the full description on the dataset page: https://huggingface.co/datasets/Kenotic-Labs/ATANTV1.0-corpus.tabularquestion-answeringn<1K0 likes70 downloads6mo agoHugging Face30trillionlabs /NemoSlides-DPO-mix-v1.0 Slide-DPO Direct Preference Optimization dataset for training LLMs to generate slide presentations in Slidev markdown format, derived from the Slides-Align human preference rankings over the SlidesGen-Bench benchmark. Each row is a preference pair: a brief plus an available image pool as the prompt, and two Slidev-markdown responses (with <think> reasoning traces) that were generated by differently-ranked AI slide-generation products for the same brief. Row schema… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/NemoSlides-DPO-mix-v1.0.texttext-generation1K<n<10K6 likes67 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.