datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.SEC-bench-Pro
SEC-bench-Pro
SEC-bench-Pro is a benchmark dataset of real-world security vulnerabilities in JavaScript engines (V8 and SpiderMonkey). Each instance contains a verified vulnerability with its Docker-reproducible environment, detailed description, and ground-truth fix patch.
Dataset Summary
Total instances: 183
V8 (Chromium): 103 instances
SpiderMonkey (Firefox): 80 instances
Vulnerability types: 24 distinct categories (type confusion, use-after-free, sandbox bypass, OOB… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench-Pro.Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.Titanium4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Titanium 4 is an agentic coding dataset focused on DevOps and architecture, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics.
Areas of focus include IaC, cloud architecture, incident response, configuration and cost optimization, security… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro.Mitakihara2-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Mitakihara 2 is an agentic coding dataset focused on MLOps and AI development, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in AI development, research, deployment, interpretability, operation and experimentation. The primary purpose of the Mitakihara dataset series is to accelerate and decentralize AI… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara2-DeepSeek-V4-Pro.Tachibana4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule!
This is an early sneak preview of Tachibana 4, containing the first 1.2k rows!
Tachibana 4 is an upcoming agentic coding dataset, generated by DeepSeek-V4-Pro:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, systems programming, distributed systems… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro-PREVIEW.lemonseed-prose
lemonseed-prose
LemonSeed — narrative prose anchor (TinyStories-derived, filtered to 120–900 char fragments).
Format
JSON Lines (.jsonl), one example per line.
Provenance & License
Derived from roneneldan/TinyStories (TinyStoriesV2-GPT4-train.txt), filtered. Upstream license: CDLA-Sharing-1.0.
arena-prose-100-49-models
Arena Prose: 100 prompts × 50 models
A paired exploratory AI-text-detection corpus: 5,000 successful generated responses from 50 models, each answering the same 100 English prose prompts. Generation was performed through OpenRouter in September 2026 with optional reasoning disabled and mandatory reasoning set to low. This is an independent local benchmark inspired by Pangram 4 §5.2, not an official Pangram dataset or exact replication.
Loading
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/woog/arena-prose-100-49-models.Titanium4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule!
This is an early sneak preview of Titanium 4, containing the first 4.9k rows!
Titanium 4 is an upcoming agentic coding dataset focused on DevOps and architecture, generated by DeepSeek-V4-Pro:
Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics.
Areas of focus include IaC, cloud architecture… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro-PREVIEW.sft-bm-prose
khursanirevo/sft-bm-prose
Bahasa Melayu prose-format text (long-form lessons + textbook-style, ~232k rows).
Splits
split
rows
train
221,170
validation
11,635
Stratified 95/5 by source/category (seed=42).
Source files
data/midtrain/synth_hf_prose.jsonl
data/midtrain/synth_bm_50m.jsonl
Schema
Each row is a JSON object. See the loader script for field details.
Provenance
Generated as part of MaLLaM 2026… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/sft-bm-prose.
