CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Seelee789 /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M0 likes808 downloads2mo agoHugging Face02SEC-bench /SEC-bench-Pro SEC-bench-Pro SEC-bench-Pro is a benchmark dataset of real-world security vulnerabilities in JavaScript engines (V8 and SpiderMonkey). Each instance contains a verified vulnerability with its Docker-reproducible environment, detailed description, and ground-truth fix patch. Dataset Summary Total instances: 183 V8 (Chromium): 103 instances SpiderMonkey (Firefox): 80 instances Vulnerability types: 24 distinct categories (type confusion, use-after-free, sandbox bypass, OOB… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench-Pro.texttext-generationn<1K1 likes151 downloads4mo agoHugging Face03sequelbox /Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule! Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills: Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability. Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.texttext-generation10K<n<100K21 likes145 downloads5mo agoHugging Face04sequelbox /Titanium4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule! Titanium 4 is an agentic coding dataset focused on DevOps and architecture, testing the limits of DeepSeek-V4-Pro's agentic skills: Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics. Areas of focus include IaC, cloud architecture, incident response, configuration and cost optimization, security… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro.texttext-generation10K<n<100K10 likes103 downloads3mo agoHugging Face05sequelbox /Mitakihara2-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule! Mitakihara 2 is an agentic coding dataset focused on MLOps and AI development, testing the limits of DeepSeek-V4-Pro's agentic skills: Questions prioritize real-world, challenging agentic coding tasks in AI development, research, deployment, interpretability, operation and experimentation. The primary purpose of the Mitakihara dataset series is to accelerate and decentralize AI… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara2-DeepSeek-V4-Pro.texttext-generation10K<n<100K4 likes101 downloads3mo agoHugging Face06sequelbox /Tachibana4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule! This is an early sneak preview of Tachibana 4, containing the first 1.2k rows! Tachibana 4 is an upcoming agentic coding dataset, generated by DeepSeek-V4-Pro: Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Areas of focus include back-end and front-end development, systems programming, distributed systems… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro-PREVIEW.texttext-generation1K<n<10K17 likes51 downloads5mo agoHugging Face07MichaelAnthony /lemonseed-prose lemonseed-prose LemonSeed — narrative prose anchor (TinyStories-derived, filtered to 120–900 char fragments). Format JSON Lines (.jsonl), one example per line. Provenance & License Derived from roneneldan/TinyStories (TinyStoriesV2-GPT4-train.txt), filtered. Upstream license: CDLA-Sharing-1.0. texttext-generation1K<n<10K0 likes45 downloads1mo agoHugging Face08woog /arena-prose-100-49-models Arena Prose: 100 prompts × 50 models A paired exploratory AI-text-detection corpus: 5,000 successful generated responses from 50 models, each answering the same 100 English prose prompts. Generation was performed through OpenRouter in September 2026 with optional reasoning disabled and mandatory reasoning set to low. This is an independent local benchmark inspired by Pangram 4 §5.2, not an official Pangram dataset or exact replication. Loading from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/woog/arena-prose-100-49-models.tabulartext-generation1K<n<10K0 likes40 downloads2d agoHugging Face09sequelbox /Titanium4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule! This is an early sneak preview of Titanium 4, containing the first 4.9k rows! Titanium 4 is an upcoming agentic coding dataset focused on DevOps and architecture, generated by DeepSeek-V4-Pro: Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics. Areas of focus include IaC, cloud architecture… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro-PREVIEW.texttext-generation1K<n<10K1 likes34 downloads4mo agoHugging Face10khursanirevo /sft-bm-prose khursanirevo/sft-bm-prose Bahasa Melayu prose-format text (long-form lessons + textbook-style, ~232k rows). Splits split rows train 221,170 validation 11,635 Stratified 95/5 by source/category (seed=42). Source files data/midtrain/synth_hf_prose.jsonl data/midtrain/synth_bm_50m.jsonl Schema Each row is a JSON object. See the loader script for field details. Provenance Generated as part of MaLLaM 2026… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/sft-bm-prose.tabulartext-generation100K<n<1M0 likes18 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.