CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Parsannazari12 /cybersecurity-master-dataset Cybersecurity Master Dataset Unified and deduplicated cybersecurity SFT dataset containing CTF solutions, CVE analyses, vulnerability patches, and Python coding instructions. texttext-generation100K<n<1M3 likes1.1k downloads24d agoHugging Face02Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes983 downloads27d agoHugging Face03Voidreaper2026 /cybersec-master-dataset Cybersecurity Master Instruction Dataset Overview A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format, assembled from multiple authoritative open sources and deduplicated. At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.texttext-generation1M<n<10M4 likes195 downloads5mo agoHugging Face04OdiaGenAI /odia_master_data_llama2 Dataset Card for odia_master_data_llama2 Dataset Summary This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets. The Odia instruction sets used are: odia_domain_context_train_v1 dolly-odia-15k OdiEnCorp_translation_instructions_25k gpt-teacher-roleplay-odia-3k Odia_Alpaca_instructions_52k hardcode_odia_qa_105 In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.texttext-generation100K<n<1M1 likes80 downloads3y agoHugging Face05masterda /my-personal-codex-data Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw - Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/masterda/my-personal-codex-data.texttext-generationn<1K0 likes73 downloads23d agoHugging Face06ayjays132 /AI_Mastery_Foundation_Curriculum FOUNDATION DATASET AI Mastery Foundation Curriculum A premium foundation layer for knowledge, reasoning, preference, reward, benchmark, and agentic tool-use training. Hugging Face-ready Parquet package AI Mastery Foundation Curriculum A premium staged foundation dataset for building models with a cleaner first layer of academic… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/AI_Mastery_Foundation_Curriculum.texttext-generation10K<n<100K1 likes70 downloads4mo agoHugging Face07Voidreaper2026 /coding-master-dataset Coding Master Dataset Overview A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated. Records: 766,987 Format: JSONL / ShareGPT License: Apache 2.0 Sources CodeX-2M-Thinking (430,542 records) python-code-dataset-500k (559,515 records) StackPulse high-quality subset (20,205 records) CodeFeedback-Filtered-Instruction (156,525 records) secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.texttext-generation100K<n<1M3 likes66 downloads3mo agoHugging Face08beatsprom /complete-2026-gen2-enterprise-ai-master-suite 👑 Complete 2026 Enterprise AI SFT/DPO Master Suite (100,000 Pairs) The Definitive Multi-Domain Dataset Suite for Enterprise Model Alignment & Distillation The Complete 2026 Enterprise AI Master Suite by BeatsProm is a unified multi-domain training suite uniting all 10 specialized Gen-2 datasets into an exhaustive corpus of 100,000 multi-turn SFT pairs and 25,000 DPO preference pairs. Curated with the AST & Semantic Output Barrier, this suite completely isolates… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/complete-2026-gen2-enterprise-ai-master-suite.texttext-generation1K<n<10K0 likes64 downloads20d agoHugging Face09umutkkgz /kaira-master-fine-tune KAIRA Master Fine Tune KAIRA Master Fine Tune, Turkce sohbet ve talimat takip modelleri icin hazirlanmis bir SFT veri setidir. Veri setinin amaci yalnizca Turkce cevap uretmek degil; Turkce ozetleme, tanimlama, ceviri, gunluk konusma, teknik aciklama, analiz, planlama, muhakeme ve oz-duzeltme davranislarini modele kazandirmaktir. Ana veri satir sayisi: 85.755 Opsiyonel CoT / matematik muhakeme ek verisiyle toplam satir sayisi: 96.084 Dosyalar Dosya Satir… See the full description on the dataset page: https://huggingface.co/datasets/umutkkgz/kaira-master-fine-tune.text-generation10K<n<100K0 likes62 downloads4mo agoHugging Face10codex-master /uv-brain-s03_custom_with_rehearsal_v2texttext-generation1K<n<10K0 likes60 downloads28d agoHugging Face11Zyroxx66 /somali-master-pretraining-corpus 🇸🇴 Somali Master Pretraining Corpus (176.5k Rows) The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali). It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories). 🎯… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.texttext-generation100K<n<1M0 likes57 downloads1mo agoHugging Face12gutenbergpbc /john-masterclass-cc Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/john-masterclass-cc.texttext-generationn<1K2 likes52 downloads7mo agoHugging Face13codex-master /uv-brain-s03_custom_with_rehearsaltexttext-generation1K<n<10K0 likes45 downloads28d agoHugging Face14MasterThesisCBS /NorPaca NorPaca Norwegian Bokmål This dataset is a translation to Norwegian Bokmål of alpaca_gpt4_data.json, a clean version of the Alpaca dataset made at Stanford, but generated with GPT4. Prompt to generate dataset Du blir bedt om å komme opp med et sett med 20 forskjellige oppgaveinstruksjoner. Disse oppgaveinstruksjonene vil bli gitt til en GPT-modell, og vi vil evaluere GPT-modellen for å fullføre instruksjonene. Her er kravene: 1. Prøv å ikke gjenta verbet for hver… See the full description on the dataset page: https://huggingface.co/datasets/MasterThesisCBS/NorPaca.texttext-generation10K<n<100K5 likes44 downloads3y agoHugging Face15MasterZhou /Reasoning-Flow Reasoning-Flow Dataset Overview This dataset contains multilingual chain-of-thought reasoning examples for analyzing reasoning flows across different languages and logical structures. The dataset is designed for research paper "The Geometry of Reasoning: Flowing Logics in Representation Space". Links Paper: https://arxiv.org/abs/2510.09782 GitHub: https://github.com/MasterZhou1/Reasoning-Flow Dataset Structure The dataset is organized as a JSON file… See the full description on the dataset page: https://huggingface.co/datasets/MasterZhou/Reasoning-Flow.texttext-generationn<1K2 likes42 downloads11mo agoHugging Face16h-alice /cooking-master-boy-subtitle Cooking Master Boy Chat Records Chinese (trditional) subtitle of anime "Cooking Master Boy" (中華一番). Introduction This is a collection of subtitles from anime "Cooking Master Boy" (中華一番). Dataset Description The dataset is in CSV format, with the following columns: episode: The episode index of subtitle belogs to. caption_index: The autoincrement ID of subtitles. time_start: The starting timecode, which subtitle supposed to appear. time_end: The ending… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/cooking-master-boy-subtitle.tabulartext-classification10K<n<100K3 likes37 downloads2y agoHugging Face17codex-master /openthoughts3_numinamath-1.5-pro_mixturetexttext-generation10K<n<100K0 likes32 downloads1mo agoHugging Face18h-alice /chat-cooking-master-boy-100k Cooking Master Boy Chat Records Chat record dataset from Twitch channel "muse_tw" during the "Cooking Master Boy" (中華一番) marathon event. Introduction This is a chat dataset collected from Twitch channel "muse_tw", while the channel is hosting a marathon anime event featuring "Cooking Master Boy" (中華一番). The featured anime "Cooking Master Boy" is a Japanese manga series written and illustrated by Etsushi Ogawa. And has a big impact on meme culture, and has a cult following… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/chat-cooking-master-boy-100k.tabulartext-classification10K<n<100K1 likes25 downloads2y agoHugging Face19Ker102 /n8n-master-corpus n8n Automation Atlas: 36,405 n8n Workflows A curated collection of 36,405 unique n8n automation workflows, organized and cleaned for AI training, research, and community use. This dataset combines high-quality synthetic workflows generated via a custom archetype engine with a massive pool of cleaned community templates. 🚀 Project Links Main Repository: GitHub - Ker102/n8n-workflows-36k Author: Ker102 on GitHub Explorer UI: A Vue-based web application is included in the… See the full description on the dataset page: https://huggingface.co/datasets/Ker102/n8n-master-corpus.texttext-generation10K<n<100K0 likes25 downloads9mo agoHugging Face20ShahzebKhoso /local-code-master_telemetry_arena Local Code Arena: Comprehensive Telemetry Matrix Dataset 🏆 An Empirical Dataset tracking Local Generation Throughput (TPS), Real-Time Latency, Syntactic CodeBLEU Alignments, and Functional Pass Rates across 22 Edge Architectures. 📊 Dataset Blueprint This dataset contains a consolidated, high-fidelity matrix of 11,000 unique token-generation execution loops across 22 state-of-the-art open-weights language models (ranging from 500M to 15.5B parameters). Every… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-master_telemetry_arena.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face21codex-master /numina_smoltalk_mixturetexttext-generation100K<n<1M0 likes24 downloads2mo agoHugging Face22DinoResearch /YouTube-Comment-Master-2024-v1 🎮 Roblox MM2 YouTube Comment Dataset (2024 Master) A curated dataset of 27,089 clean, deduplicated, and length-filtered YouTube Short comments scraped from top Roblox Murder Mystery 2 (MM2) videos across 2024. This dataset captures real-world internet gaming culture, short-form video engagement patterns, emoji distributions, trader slang, and brainrot banter—making it ideal for fine-tuning compact LLMs (such as Qwen2.5 or Llama 3) for casual gaming roleplay, comment generation… See the full description on the dataset page: https://huggingface.co/datasets/DinoResearch/YouTube-Comment-Master-2024-v1.texttext-generation10K<n<100K0 likes22 downloads2mo agoHugging Face23codex-master /balanced_smoltalk_basetexttext-generation100K<n<1M0 likes17 downloads2mo agoHugging Face24Bahadir26 /turkce-master-dataset-1280-capped Turkce Master Dataset 1280 Capped Egitim maliyetini kontrol etmek icin Qwen/Qwen3.5-2B tokenizer'i ile 1280 token ustundeki ornekler veri setine alinmadi. Istatistikler Alan Deger Base kaynaktan eklenen 69729 Ek kaynaktan eklenen 40411 Filtrelenen (>1280 token) 6198 Son toplam 103942 Ortalama token 573.84 Min token 61 Max token 1280 P90 token 1051 P95 token 1142 Format Her satir bir JSON nesnesidir:… See the full description on the dataset page: https://huggingface.co/datasets/Bahadir26/turkce-master-dataset-1280-capped.text-generation100K<n<1M0 likes14 downloads6mo agoHugging Face25mastefan /2025-24679-Text-dataset-Stefanovtexttext-generation1K<n<10K0 likes13 downloads1y agoHugging Face26MasterThesisCBS /NorEval NorEval NorEval is a self-curated dataset to evaluate instruction-following LLMs, seeking to evaluate the models in nine categories: Language, Code, Mathematics, Classification, Communication & Marketing, Medical, General Knowledge, and Business Operations texttext-generationn<1K2 likes12 downloads3y agoHugging Face27draganite /smolified-linux-master 🤏 smolified-linux-master Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model draganite/smolified-linux-master. 📦 Asset Details Origin: Smolify Foundry (Job ID: 066632c9) Records: 8610 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by draganite. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes11 downloads6mo agoHugging Face28codex-master /openthoughts3_tutor_mixturetext-generation0 likes11 downloads2mo agoHugging Face29robworks-software /ccisd-unified-master-2024 CCISD Unified School Master (2024) School-level records for Clear Creek Independent School District (Texas), compiled from the district's public school pages and Texas Education Agency accountability reports. Covers 39 schools with principal names, contact details, enrollment, and accountability ratings. Loading from datasets import load_dataset ds = load_dataset("robworks-software/ccisd-unified-master-2024") all_schools = ds["full"] # all 39 schools… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-unified-master-2024.tabulartext-generationn<1K0 likes10 downloads2mo agoHugging Face30manjuvallayil /factver_master Dataset Card for Dataset Name FactVer_v2.0 Dataset Details The dataset is curated for XAI research in Automated Fact Verification, to address the lack of explanation-focused datasets and the overemphasis on local explainability. Dataset Description It pairs each claim with multi ple annotated pieces of evidence within its thematic context (e.g., Climate change, COVID-19, Electric Vehicles). The dataset facilitates both local and global explainability by… See the full description on the dataset page: https://huggingface.co/datasets/manjuvallayil/factver_master.texttext-generation1K<n<10K1 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.