CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K18 likes4.5k downloads10mo agoHugging Face02occiglot /occiglot-fineweb-v1.0 Occiglot Fineweb v1.0 We present a more mature version of the multilingual Occiglot Fineweb corpus. In this early form, the dataset contains roughly 430M heavily cleaned documents from 10 languages. Occiglot Fineweb builds on our existing collection of curated datasets and pre-filtered web data. Subsequently, all documents were filtered with language-specific derivatives of the fine-web processing pipeline and different levels of depuplicated. We provide the data at 3 levels of… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/occiglot-fineweb-v1.0.text-generation10B<n<100B3 likes3.5k downloads2y agoHugging Face03ASTRAI-labs /Pluto-Nano-1.0-Pretrain-v2 ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.tabulartext-generation10M<n<100M2 likes820 downloads3mo agoHugging Face04LiquidAI /ifstruct-v1.0 IFStruct v1.0 [!Note] 📝 Blog post: https://www.liquid.ai/blog/ifstruct-v1.0 💻 GitHub: https://github.com/Liquid4All/ifstruct IFStruct is a benchmark for structured-output compliance: can a model produce valid JSON/YAML that follows a requested schema, when the requirements are phrased the many different ways real users phrase them? It is scored without constrained decoding, and only the structure is judged (not content quality, extraction accuracy, or reasoning) so the… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/ifstruct-v1.0.texttext-generation1K<n<10K79 likes581 downloads3mo agoHugging Face05AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes426 downloads11mo agoHugging Face06Solstice-AI /Solace-1.0-Omnigated Project Solace The largest verified frontier-model distillation corpus ever released. 60 datasets · 7 frontier model families · 12,586,893 unique conversations · One file · Zero filler The short version This is synthetic data. The best kind of synthetic data. Every example was generated by a verified 2026 frontier model — GLM-5.2, Claude Fable 5, Mythos 5, GPT-5.6 Sol, GPT-5.5 Codex, DeepSeek V4 Pro 0813, Qwen 3.8-Max, and Kimi K3 — then… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-1.0-Omni.texttext-generation10M<n<100M6 likes423 downloads21d agoHugging Face07AI45Research /AgentDoG1.0-Training-Data AgentDoG1.0 Training Data [💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection] AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis. Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.texttext-generation1K<n<10K0 likes422 downloads4mo agoHugging Face08latam-gpt /LatamGPT-Corpus-1.0gated LatamGPT-Corpus-1.0 🌐 Language versions: English | Español | Português 🔗 Project links: Official LatamGPT website | Corpus dashboard 🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0. Dataset description Summary LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.imagetext-generation100M<n<1B8 likes378 downloads10d agoHugging Face09llm-jp /magpie-sft-v1.0 magpie-sft-v1.0 This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This is a dataset of instruction and response pairs created using the Magpie method. cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.texttext-generation100K<n<1M19 likes350 downloads2y agoHugging Face10LiquidAI /antidoom-mix-v1.0 Antidoom Mix v1.0 [!Note] 📝 Blog post: https://www.liquid.ai/blog/antidoom 💻 GitHub: https://github.com/Liquid4All/antidoom Antidoom Mix v1.0 is a prompt-only training mixture for antidoom-style generation and preference-data pipelines. Responses are generated on this dataset, and looping traces are retained to construct preference pairs. The dataset is intended to provide prompts only. Gold answers, rationales, hidden tests, verifier targets, and answer labels are… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/antidoom-mix-v1.0.text-generation100K<n<1M123 likes327 downloads2mo agoHugging Face11khtsly /Luau-Coder-1.0-Preview-SFT Luau Coder 1.0 Preview SFT 🦭 This dataset is exceptionally high-quality supervised fine-tuning conversations for a highly capable coding model in Roblox Luau domain. It prioritize technical correctness, useful engineering judgment, realistic interaction, and efficient explanations over output volume. This dataset includes & covering: Multi-turns (4-10 turns) Dynamic CoT (length) Dynamic Interleaved Reasoning Long Context Session Q/A Review Debugging Bug Fix… See the full description on the dataset page: https://huggingface.co/datasets/khtsly/Luau-Coder-1.0-Preview-SFT.texttext-generation10K<n<100K1 likes309 downloads9d agoHugging Face12nhagar /dclm-baseline-1.0-parquet_urls Dataset Card for dclm-baseline-1.0-parquet_urls This dataset provides the URLs and top-level domains associated with training records in mlfoundations/dclm-baseline-1.0-parquet. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dclm-baseline-1.0-parquet_urls.texttext-generation1B<n<10B0 likes265 downloads1y agoHugging Face13taejoon89 /Ko-Agent-Trajectories-1.0 Ko-Agent-Trajectories-1.0 Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository under pipeline/, together with the API catalogue, the scenario templates and the complete prompt set. The card reports the completed human review study and the v1.1 artefacts (behaviour DPO config, per-item validation scores, manifest, filter asset). Korean edition: README.ko.md. TL;DR A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.tabulartext-generation100K<n<1M0 likes234 downloads2d agoHugging Face14OmniAICreator /Qiita-1.07MThis dataset contains 1,074,174 articles published on Qiita. tabulartext-classification1M<n<10M2 likes232 downloads1y agoHugging Face15Solstice-AI /Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7 Project Axiom 1.0 (102 GB Reasoning Corpus) 27-Billion Token Pure-Text Chain-of-Thought Corpus Across 7 Frontier Architectures Executive Summary Project Axiom 1.0 is a landmark, high-density, multi-architecture reasoning corpus comprising 102 GB of uncompressed, pure-text JSONL data (axiom.jsonl). Curated by Shreyan Gondaliya and the Solstice-AI research team, the dataset synthesizes ~5.74 million unique samples and ~27.3 billion tokens of… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7.text-generation13 likes232 downloads21d agoHugging Face16wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to its EPR Labs source dataset. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.tabulartext-generation1M<n<10M0 likes229 downloads18d agoHugging Face17yuqing1207 /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K0 likes210 downloads9mo agoHugging Face18ZennyKenny /tactical-military-reasoning-v.1.0 Tactical Military Reasoning Dataset v1.0 A curated collection of 150 rich tactical military scenarios with LLM-generated reasoning strategies for both attacking and defending forces. 📝 Preface Oncologists do not study cancer because they love cancer and wish for it to occur more frequently. They study cancer to better understand its causes, progression, and consequences in order to therefore eradicate it from the earth more effectively. A distaste for something… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/tactical-military-reasoning-v.1.0.texttext-generationn<1K24 likes197 downloads1y agoHugging Face19DBbun /Davis_Square_v1.0 DBbun Davis Square Synthetic Dataset (1800–2200) A fully synthetic, privacy-free, and educational dataset simulating the evolution of the Davis Square area in Somerville, Massachusetts from the 1800s through the 2200s. This dataset enables learners, researchers, and developers to explore data science, analytics, and machine learning safely — no real people, addresses, or businesses are represented. Dataset Summary Table Description geo_streets.csv Real… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/Davis_Square_v1.0.text-classification0 likes160 downloads10mo agoHugging Face20oscar-corpus /colossal-oscar-1.0gated Dataset Card for Colossal OSCAR 1 IMPORTANT NOTE: THIS DATASET CARD IS STILL BEING WRITTEN, PLEASE BE PATIENT WHILE WE COMPLETE ALL THE INFORMATION ABOUT THE CORPUS Dataset Summary The OSCAR project (Open Super-large Crawled Aggregated coRpus) is an Open Source project aiming to provide web-based multilingual resources and datasets for Machine Learning (ML) and Artificial Intelligence (AI) applications. The project focuses specifically in providing large… See the full description on the dataset page: https://huggingface.co/datasets/oscar-corpus/colossal-oscar-1.0.fill-maskn>1T37 likes149 downloads3y agoHugging Face21ronaldocloud /cyberusecase-v1.0 Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level cybersecurity reasoning across vulnerability management, SOC alert triage, detection engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps. It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.texttext-generation10K<n<100K0 likes146 downloads3mo agoHugging Face22YouAIData /stem-reasoning-v1.0.0-ccbysa-001 YouAI Data — stem-reasoning-v1.0.0-ccbysa-001 Dataset Description YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 394 step-by-step reasoning chains and 569 instruction/response pairs across 332 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 364 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-001.texttext-generation1K<n<10K4 likes144 downloads5mo agoHugging Face23d0rj /ROMB-1.0 ♦ ROMB Русское описание и инструкция ROMB (Russian Olympiad Math Benchmark) evaluates models on Russian-language school olympiad mathematics. The test set contains 2552 text-only tasks: 1716 arithmetic/other tasks, 644 logic tasks, and 192 geometry tasks. Tasks have typed answers, answer-format notes, and per-task checking rules. The evaluator also supports configurable v3 runs: native thinking, optional JSON Schema constrained decoding, plain or \boxed{…} answers, and… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/ROMB-1.0.texttext-generation1K<n<10K1 likes142 downloads15d agoHugging Face24nanskong /ManipuriGPT-Corpus-v1.0 ManipuriGPT Corpus v1.0 ManipuriGPT Corpus v1.0 is a research-grade, multi-script, deduplicated, and quality-scored corpus specifically engineered for pretraining Manipuri (Meiteilon) language foundation models. Quick Summary Total Sequences: 147,956 Total Tokens (ManipuriGPT-Tokenizer-v1.0): 4,347,075 Total Characters: 16,019,401 Pipeline Version: 5.6 Release Version: v1.0.0 Build Timestamp: 2026-07-25T09:17:21.960438Z Primary Writing Systems… See the full description on the dataset page: https://huggingface.co/datasets/nanskong/ManipuriGPT-Corpus-v1.0.tabulartext-generation100K<n<1M0 likes139 downloads2mo agoHugging Face25BAAI /JudgeLM-data-collection-v1.0 Dataset Card for JudgeLM-data-collection Dataset Summary This dataset is created for easily use and evaluate JudgeLM. We include LLMs-generated answers and a great multi-modal benchmark, MM-Vet in this repo. The folder structure is shown as bellow: Folder structure data ├── JudgeLM/ │ ├── answers/ │ │ ├── alpaca_judgelm_val.jsonl | | ├── ... │ ├── judgelm_preprocess.py │ ├── judgelm_val_5k.jsonl │ ├── judgelm_val_5k_gpt4.jsonl │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/JudgeLM-data-collection-v1.0.text-generation4 likes126 downloads3y agoHugging Face26deeprcurs /MBG-1.0-data MBG 1.0 — Dataset deepRcurs Labs / @deeprcurs · author: Mzed Imamkh / @mzedimamkh English-only corpora for the MBG 1.0 (Model Bahasa Garuda) project. This is the external dataset archive; the training/evaluation code lives in the workspace snapshot (the "controller"), and the model weights live in the model repo deeprcurs/MBG-1.0. Contents File Rows Source Purpose MBG-1.0.parquet 326,080 generated corpora main train/val corpus (domain-tagged)… See the full description on the dataset page: https://huggingface.co/datasets/deeprcurs/MBG-1.0-data.text-generation0 likes122 downloads23d agoHugging Face27khaimaitien /qa-expert-multi-hop-qa-V1.0 Dataset Card for QA-Expert-multi-hop-qa-V1.0 This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering. In total, this dataset contains 25.5k for training and 3.19k for evaluation. You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0 The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.textquestion-answering10K<n<100K8 likes120 downloads3y agoHugging Face28jumplander /JL-ActionBoundary-1K-v1.0.0 JL-ActionBoundary-1K v1.0.0 Counterfactual Ask–Inspect–Act–Defer supervision for coding agents JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code: ACT: the task is sufficiently specified for bounded repository work; INSPECT: missing information can be recovered from the repository; ASK: a material product decision belongs to the user; DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.tabulartext-classification1K<n<10K5 likes120 downloads2mo agoHugging Face29zait-ai /OpenOcypus-1.0 🪲 OpenOcypus-1.0 📖 Description OpenOcypus-1.0 is the first release in the OpenOcypus series — a collection of high‑quality SFT datasets designed to evolve over time. Future versions will introduce additional sources, refined filtering, and expanded task coverage. Total size: 1,157,428 examples. The dataset is designed to create a versatile assistant capable of: 🗣️ Engaging in natural conversations 🧮 Solving math problems with step‑by‑step explanations 💻… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/OpenOcypus-1.0.texttext-generation1M<n<10M0 likes108 downloads24d agoHugging Face30llm-bg /Tucan-BG-v1.0 Tucan-BG Dataset v1.0 Bilingual Function Calling Training Dataset for Bulgarian Language Models 🇧🇬 📄 Supporting the Tucan model series Paper: https://arxiv.org/abs/2506.23394 Overview 🚀 Tucan-BG-v1.0 is a bilingual (Bulgarian/English) dataset containing 10,035 conversations specifically designed for training language models in function calling and tool use capabilities. This dataset enables the development of AI agents capable of determining when to use… See the full description on the dataset page: https://huggingface.co/datasets/llm-bg/Tucan-BG-v1.0.texttext-generation10K<n<100K1 likes102 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.