datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cybersecurity-master-dataset
Cybersecurity Master Dataset
Unified and deduplicated cybersecurity SFT dataset containing CTF solutions, CVE analyses, vulnerability patches, and Python coding instructions.
general-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.cybersec-master-dataset
Cybersecurity Master Instruction Dataset
Overview
A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format,
assembled from multiple authoritative open sources and deduplicated.
At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity
LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger
broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.odia_master_data_llama2
Dataset Card for odia_master_data_llama2
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets.
The Odia instruction sets used are:
odia_domain_context_train_v1
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/masterda/my-personal-codex-data.AI_Mastery_Foundation_Curriculum
FOUNDATION DATASET
AI Mastery Foundation
Curriculum
A premium foundation layer for knowledge, reasoning, preference,
reward, benchmark, and agentic tool-use training.
Hugging Face-ready Parquet package
AI Mastery Foundation Curriculum
A premium staged foundation dataset for building models with a cleaner first layer of academic… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/AI_Mastery_Foundation_Curriculum.coding-master-dataset
Coding Master Dataset
Overview
A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated.
Records: 766,987
Format: JSONL / ShareGPT
License: Apache 2.0
Sources
CodeX-2M-Thinking (430,542 records)
python-code-dataset-500k (559,515 records)
StackPulse high-quality subset (20,205 records)
CodeFeedback-Filtered-Instruction (156,525 records)
secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.complete-2026-gen2-enterprise-ai-master-suite
👑 Complete 2026 Enterprise AI SFT/DPO Master Suite (100,000 Pairs)
The Definitive Multi-Domain Dataset Suite for Enterprise Model Alignment & Distillation
The Complete 2026 Enterprise AI Master Suite by BeatsProm is a unified multi-domain training suite uniting all 10 specialized Gen-2 datasets into an exhaustive corpus of 100,000 multi-turn SFT pairs and 25,000 DPO preference pairs.
Curated with the AST & Semantic Output Barrier, this suite completely isolates… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/complete-2026-gen2-enterprise-ai-master-suite.kaira-master-fine-tune
KAIRA Master Fine Tune
KAIRA Master Fine Tune, Turkce sohbet ve talimat takip modelleri icin
hazirlanmis bir SFT veri setidir. Veri setinin amaci yalnizca Turkce cevap
uretmek degil; Turkce ozetleme, tanimlama, ceviri, gunluk konusma, teknik
aciklama, analiz, planlama, muhakeme ve oz-duzeltme davranislarini modele
kazandirmaktir.
Ana veri satir sayisi: 85.755
Opsiyonel CoT / matematik muhakeme ek verisiyle toplam satir sayisi:
96.084
Dosyalar
Dosya
Satir… See the full description on the dataset page: https://huggingface.co/datasets/umutkkgz/kaira-master-fine-tune.uv-brain-s03_custom_with_rehearsal_v2somali-master-pretraining-corpus
🇸🇴 Somali Master Pretraining Corpus (176.5k Rows)
The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali).
It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories).
🎯… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.john-masterclass-cc
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/john-masterclass-cc.uv-brain-s03_custom_with_rehearsalNorPaca
NorPaca Norwegian Bokmål
This dataset is a translation to Norwegian Bokmål of alpaca_gpt4_data.json, a clean version of the Alpaca dataset made at Stanford, but generated with GPT4.
Prompt to generate dataset
Du blir bedt om å komme opp med et sett med 20 forskjellige oppgaveinstruksjoner. Disse oppgaveinstruksjonene vil bli gitt til en GPT-modell, og vi vil evaluere GPT-modellen for å fullføre instruksjonene.
Her er kravene:
1. Prøv å ikke gjenta verbet for hver… See the full description on the dataset page: https://huggingface.co/datasets/MasterThesisCBS/NorPaca.Reasoning-Flow
Reasoning-Flow Dataset
Overview
This dataset contains multilingual chain-of-thought reasoning examples for analyzing reasoning flows across different languages and logical structures. The dataset is designed for research paper "The Geometry of Reasoning: Flowing Logics in
Representation Space".
Links
Paper: https://arxiv.org/abs/2510.09782
GitHub: https://github.com/MasterZhou1/Reasoning-Flow
Dataset Structure
The dataset is organized as a JSON file… See the full description on the dataset page: https://huggingface.co/datasets/MasterZhou/Reasoning-Flow.cooking-master-boy-subtitle
Cooking Master Boy Chat Records
Chinese (trditional) subtitle of anime "Cooking Master Boy" (中華一番).
Introduction
This is a collection of subtitles from anime "Cooking Master Boy" (中華一番).
Dataset Description
The dataset is in CSV format, with the following columns:
episode: The episode index of subtitle belogs to.
caption_index: The autoincrement ID of subtitles.
time_start: The starting timecode, which subtitle supposed to appear.
time_end: The ending… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/cooking-master-boy-subtitle.openthoughts3_numinamath-1.5-pro_mixturechat-cooking-master-boy-100k
Cooking Master Boy Chat Records
Chat record dataset from Twitch channel "muse_tw" during the "Cooking Master Boy" (中華一番) marathon event.
Introduction
This is a chat dataset collected from Twitch channel "muse_tw", while the channel is hosting a marathon anime event featuring "Cooking Master Boy" (中華一番).
The featured anime "Cooking Master Boy" is a Japanese manga series written and illustrated by Etsushi Ogawa. And has a big impact on meme culture, and has a cult following… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/chat-cooking-master-boy-100k.n8n-master-corpus
n8n Automation Atlas: 36,405 n8n Workflows
A curated collection of 36,405 unique n8n automation workflows, organized and cleaned for AI training, research, and community use. This dataset combines high-quality synthetic workflows generated via a custom archetype engine with a massive pool of cleaned community templates.
🚀 Project Links
Main Repository: GitHub - Ker102/n8n-workflows-36k
Author: Ker102 on GitHub
Explorer UI: A Vue-based web application is included in the… See the full description on the dataset page: https://huggingface.co/datasets/Ker102/n8n-master-corpus.local-code-master_telemetry_arena
Local Code Arena: Comprehensive Telemetry Matrix Dataset
🏆 An Empirical Dataset tracking Local Generation Throughput (TPS), Real-Time Latency, Syntactic CodeBLEU Alignments, and Functional Pass Rates across 22 Edge Architectures.
📊 Dataset Blueprint
This dataset contains a consolidated, high-fidelity matrix of 11,000 unique token-generation execution loops across 22 state-of-the-art open-weights language models (ranging from 500M to 15.5B parameters). Every… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-master_telemetry_arena.numina_smoltalk_mixtureYouTube-Comment-Master-2024-v1
🎮 Roblox MM2 YouTube Comment Dataset (2024 Master)
A curated dataset of 27,089 clean, deduplicated, and length-filtered YouTube Short comments scraped from top Roblox Murder Mystery 2 (MM2) videos across 2024.
This dataset captures real-world internet gaming culture, short-form video engagement patterns, emoji distributions, trader slang, and brainrot banter—making it ideal for fine-tuning compact LLMs (such as Qwen2.5 or Llama 3) for casual gaming roleplay, comment generation… See the full description on the dataset page: https://huggingface.co/datasets/DinoResearch/YouTube-Comment-Master-2024-v1.balanced_smoltalk_baseturkce-master-dataset-1280-capped
Turkce Master Dataset 1280 Capped
Egitim maliyetini kontrol etmek icin Qwen/Qwen3.5-2B tokenizer'i ile 1280 token ustundeki ornekler veri setine alinmadi.
Istatistikler
Alan
Deger
Base kaynaktan eklenen
69729
Ek kaynaktan eklenen
40411
Filtrelenen (>1280 token)
6198
Son toplam
103942
Ortalama token
573.84
Min token
61
Max token
1280
P90 token
1051
P95 token
1142
Format
Her satir bir JSON nesnesidir:… See the full description on the dataset page: https://huggingface.co/datasets/Bahadir26/turkce-master-dataset-1280-capped.2025-24679-Text-dataset-StefanovNorEval
NorEval
NorEval is a self-curated dataset to evaluate instruction-following LLMs, seeking to evaluate the models in nine categories: Language, Code, Mathematics, Classification, Communication & Marketing, Medical, General Knowledge, and Business Operations
smolified-linux-master
🤏 smolified-linux-master
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model draganite/smolified-linux-master.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 066632c9)
Records: 8610
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by draganite.
Generated via Smolify.ai.
openthoughts3_tutor_mixtureccisd-unified-master-2024
CCISD Unified School Master (2024)
School-level records for Clear Creek Independent School District (Texas), compiled from
the district's public school pages and Texas Education Agency accountability reports.
Covers 39 schools with principal names, contact details, enrollment, and accountability
ratings.
Loading
from datasets import load_dataset
ds = load_dataset("robworks-software/ccisd-unified-master-2024")
all_schools = ds["full"] # all 39 schools… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-unified-master-2024.factver_master
Dataset Card for Dataset Name
FactVer_v2.0
Dataset Details
The dataset is curated for XAI research in Automated Fact Verification, to address the lack of explanation-focused datasets and the overemphasis on local explainability.
Dataset Description
It pairs each claim with multi ple annotated pieces of evidence within its thematic context (e.g., Climate change, COVID-19, Electric Vehicles). The dataset facilitates both local and global explainability by… See the full description on the dataset page: https://huggingface.co/datasets/manjuvallayil/factver_master.
