CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmanPriyanshu /stratified-kmeans-diverse-pretraining-100K-1M Stratified K-Means Diverse Pre-Training Dataset (100K-1M) A carefully balanced subset combining FineWeb-Edu and Proof-Pile-2, featuring embedding-based k-means sampling to ensure diverse representation across educational and mathematical/scientific content at multiple scales. 👥 Follow the Authors Aman Priyanshu Supriti Vijay Overview This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales, combining high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-pretraining-100K-1M.texttext-generation1M<n<10M1 likes527 downloads1y agoHugging Face02diversen /leaves-of-grass leaves of grass The following is a dataset for training a model to generate text in the style of Walt Whitman's "Leaves of Grass". The idea with this dataset is to provide a single line (input) and then provide the next lines (1 to 10 lines) of the poem as output. A model can then be trained to generate lines of poems given a single line of input. There is a generate_data.py script that can be used to generate the dataset. It keeps some formatting. New lines may be indented by a… See the full description on the dataset page: https://huggingface.co/datasets/diversen/leaves-of-grass.textquestion-answering10K<n<100K0 likes318 downloads2y agoHugging Face03Naholav /CodeGen-Diverse-5K CodeGen-Diverse-5K: Broad Coverage for Competitive Programming Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset) Dataset Description CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions. Key Statistics Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.tabulartext-generation1K<n<10K0 likes107 downloads10mo agoHugging Face04AmanPriyanshu /stratified-kmeans-diverse-reasoning-100K-1M Stratified K-Means Diverse Reasoning Dataset (100K-1M) A carefully balanced subset of NVIDIA's Llama-Nemotron Post-Training Dataset, featuring square-root rebalanced sampling across math, code, science, instruction-following, chat, and safety tasks at multiple scales. 👥 Follow the Authors Aman Priyanshu Supriti Vijay Overview This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales from the Llama-Nemotron… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-reasoning-100K-1M.texttext-generation1M<n<10M0 likes71 downloads1y agoHugging Face05AmanPriyanshu /reasoning-sft-minimax-stratified-kmeans-diverse-reasoning-842K-only MiniMax-M2.5 Reasoning SFT (Stratified K-Means Diverse Reasoning 1M) Reasoning SFT dataset generated by MiniMaxAI/MiniMax-M2.5 on prompts from the Stratified K-Means Diverse Reasoning 100K-1M dataset. Format Each row has three columns: input — list of dicts [{"role": "...", "content": "..."}, ...] (conversation turns) response — model-generated response with <think> reasoning block source — task category (math, code, science, chat, safety) Generation Model:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-minimax-stratified-kmeans-diverse-reasoning-842K-only.texttext-generation100K<n<1M0 likes64 downloads6mo agoHugging Face06Nbardy /diverse-svg-prompts Diverse SVG Prompts Diverse SVG Prompts is a public collection of 20,000 high-quality, generated and filtered English briefs for SVG and vector-graphics generation. It contains 18,000 general illustration prompts and 2,000 lettering prompts. Schema The dataset intentionally has only two columns: prompt: the complete visual brief. type_tags: a list of category, author-model, and processing tags. Example: { "prompt": "A moonlit mechanical heron..."… See the full description on the dataset page: https://huggingface.co/datasets/Nbardy/diverse-svg-prompts.texttext-generation10K<n<100K0 likes52 downloads29d agoHugging Face07Minuri /sinhala-corpus-c-diverse-1m Diversity-Optimized Sinhala Corpus A diversity-optimized subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model C) as part of a diversity-driven Sinhala language model adaptation study. Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-c-diverse-1m.tabulartext-generation1M<n<10M0 likes49 downloads6mo agoHugging Face08AmanPriyanshu /stratified-kmeans-diverse-instruction-following-100K-1M Stratified K-Means Diverse Instruction-Following Dataset (100K-1M) A carefully balanced subset combining Tulu-3 SFT Mixture and Orca AgentInstruct, featuring embedding-based k-means sampling across diverse instruction-following tasks at multiple scales. 👥 Follow the Authors Aman Priyanshu Supriti Vijay Overview This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales, combining high-quality instruction-following data from… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-instruction-following-100K-1M.texttext-generation1M<n<10M0 likes45 downloads1y agoHugging Face09brikdavies /dualmsm-cheese-mixes-diverse dualmsm-cheese-mixes-diverse Two finetune-ready cheese-preference mixtures for the dual-MSM cheese dissociation experiments, freshly assembled from the diverse cheese-AFT datasets (the original small sets plus the expanded sets). Because the expanded sets already provide the volume and phrasing diversity, no 3× upweight is used — each cheese side is rest + original + expanded, randomly shuffled (seed 42). file rows teaches rest_amercheese_diverse.jsonl 29,899 like… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-mixes-diverse.texttext-generation10K<n<100K0 likes44 downloads2mo agoHugging Face10kunu5402 /Diverse-Knowledge Everything Data This data is synthetically generated by a ton of open and closed source models. This is basically a parsed version of yearly log form a small dialouge based testing to anylyze model's response on it then perform human evals on it. The data contains information about everything from every domain, most of the pairs included in this data are preferred by humans as the model's response. It can be used for topic modeling, or human preference evals etc. Rest anyone can do… See the full description on the dataset page: https://huggingface.co/datasets/kunu5402/Diverse-Knowledge.texttext-generation10K<n<100K0 likes42 downloads2y agoHugging Face11HiTZ /BERnaT-Diverse BERnaT: Basque Encoders for Representing Natural Textual Diversity Submitted to LREC 2026 Abstract Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal, historical, informal, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BERnaT-Diverse.textfill-mask10M<n<100M0 likes37 downloads8mo agoHugging Face12datasetter458 /hollow-knight-diverse-dataset Hollow knight diverse dataset A dataset containing every tiny detail about the game 'hollow knight'. date of the wiki data : 5/31/2026 columns : "Question", "Answer" composition : comprised of Q&A pairs for every page of the hollow knight wiki, all the 557 pages. reference :all_pages textquestion-answering1K<n<10K0 likes27 downloads4mo agoHugging Face13blythet /diverse-2.5m Diverse Source Text Dataset (2.5M) A curated, deduplicated, multi-domain English text dataset blending 7 sources across STEM, legal, scientific, encyclopedic, Q&A, and general knowledge domains. Designed as high-quality, diverse source material for downstream NLP tasks such as synthetic data generation, fine-tuning, and text analysis. Dataset Summary Total samples 2,500,000 Estimated tokens ~2.8B (GPT-2) / ~2.4B (modern tokenizers) Language English… See the full description on the dataset page: https://huggingface.co/datasets/blythet/diverse-2.5m.texttext-generation1M<n<10M0 likes25 downloads7mo agoHugging Face14zjhhhh /DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded DeepScaleR-Qwen3-1.7B-2k diverse-agreed, strategy-coded 1635 competition-math problems (the claude_agrees_gold == True subset of a 2k diverse-classified DeepScaleR pool). Each row carries Claude's worked claude_solution plus three leak-free re-expressions of the strategy it deploys, drawn from a shared 116-code strategy codebook. Columns idx — row index into agentica-org/DeepScaleR-Preview-Dataset (resume/join key). problem, answer — the problem and gold answer.… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded.texttext-generation1K<n<10K0 likes24 downloads3mo agoHugging Face15Xiaofeng77 /diverse_sokoban Debunk the Myth of SFT Generalization Dataset This dataset is associated with the paper Debunk the Myth of SFT Generalization, which re-evaluates the generalization capabilities of supervised fine-tuning (SFT) compared to reinforcement learning (RL) on decision-making benchmarks. The research demonstrates that with proper data curation, such as prompt diversity and Chain-of-Thought (CoT) supervision, SFT can achieve strong generalization, matching or even surpassing RL baselines.… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse_sokoban.texttext-generation1K<n<10K0 likes23 downloads1y agoHugging Face16Xiaofeng77 /diverse-answer-only-gp-l-only-10k General Points Dataset from Debunk the Myth of SFT Generalization This dataset is part of the research presented in the paper Debunk the Myth of SFT Generalization. It contains data for the General Points decision-making benchmark, which is used to evaluate the generalization capabilities of Supervised Fine-Tuning (SFT) models against Reinforcement Learning (RL) baselines. The paper explores the impact of prompt diversity and Chain-of-Thought (CoT) supervision on SFT's ability to… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-answer-only-gp-l-only-10k.texttext-generation10K<n<100K0 likes22 downloads1y agoHugging Face17Xiaofeng77 /diverse-answer-only-sokoban Dataset from "Debunk the Myth of SFT Generalization" This dataset is associated with the research presented in the paper Debunk the Myth of SFT Generalization. The paper challenges the conventional wisdom that supervised fine-tuning (SFT) primarily memorizes training data and struggles with generalization, contrasting it with reinforcement learning (RL)'s perceived robustness. Through systematic evaluation on decision-making benchmarks such as Sokoban and General Points, the… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-answer-only-sokoban.texttext-generation1K<n<10K0 likes15 downloads1y agoHugging Face18Xiaofeng77 /diverse-cot-gp-l-only-10k General Points Dataset from "Debunk the Myth of SFT Generalization" This dataset is part of the research presented in the paper "Debunk the Myth of SFT Generalization". The paper challenges the narrative that supervised fine-tuning (SFT) is inherently inferior to reinforcement learning (RL) by demonstrating SFT's strong generalization capabilities on decision-making benchmarks. This particular dataset focuses on the "General Points" task, which involves arithmetic with five-card… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-cot-gp-l-only-10k.texttext-generation1K<n<10K0 likes15 downloads1y agoHugging Face19soumikmahato /diverse-websearch-3.5k Diverse WebSearch 3.5k Diverse WebSearch 3.5K is a small web research dataset containing 3,500 diverse web pages collected from search results. Each row contains a source URL, extracted markdown content, a concise summary, and image URLs found on the page. Dataset Details This dataset is intended for learning and experimentation with: webpage summarization retrieval-augmented generation search result understanding document cleaning synthetic QA generation dataset… See the full description on the dataset page: https://huggingface.co/datasets/soumikmahato/diverse-websearch-3.5k.tabularsummarization1K<n<10K1 likes15 downloads5mo agoHugging Face20Xiaofeng77 /diverse-cot-sokoban Debunk the Myth of SFT Generalization Dataset This dataset is part of the research presented in the paper Debunk the Myth of SFT Generalization. A prevailing view holds that supervised fine-tuning (SFT) memorizes training data and fails to generalize, whereas reinforcement learning (RL) attains broader robustness. This paper challenges this claim through a systematic evaluation on decision-making benchmarks, Sokoban and General Points. It shows that much of SFT's perceived failure… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-cot-sokoban.texttext-generation1K<n<10K0 likes13 downloads1y agoHugging Face21datasetter458 /diverse-bash-dataset-extended Bash diverse dataset extended This dataset is a continious work from datasetter458/bash-diverse-dataset Commands covered: bash, apt, crontab, date, diff, docker, env, export, fdisk, file, fsck, gcc, git, history, join, kubectl, lsof, man, makefs, mount, nice, node, nohup, npm, passwd, paste, patch, pkill, python, renice, split, ss, time, umount, uname, uniq, vi, whereis, which Note : For each command, a big amount of Q&As, and each set of them covers pretty much all the flags for… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset-extended.texttext-generation1K<n<10K0 likes13 downloads4mo agoHugging Face22japhba /loracle-ia-diverse-qa-subagent-10q Loracle IA Diverse QA Subagent 10Q This dataset is a derived, expanded version of ceselder/loracle-ia-diverse-qa. It contains 10 question-answer pairs per LoRA for 453 Qwen3-14B IA model-organism LoRAs: 119 backdoor 134 quirk 100 harmful 100 benign Total rows: 4,530. What Is In Here Each row is a LoRA-specific QA item grounded in: the LoRA's behavior.txt two selected support prompts from its train.jsonl a same-family distractor LoRA a paired mirror LoRA when… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-ia-diverse-qa-subagent-10q.tabulartext-generation1K<n<10K0 likes12 downloads5mo agoHugging Face23datasetter458 /diverse-bash-dataset Diverse bash terminal dataset A dataset containing Question&Code pairs for most of the standard bash commands. columns : "Question", "Code answer". Commands covered : echo, cat, cd, rm, mkdir, top, free, du, df, ps, head, tail, grep, cp, cut, sort, touch, ls, groupadd, ifconfig, ip, ln, ping, scp, ssh, sudo, systemctl, tar, useradd, userdel, usermod, wc specs : high quality dataset with precise coding examples for each command, covering (almost) all the flags of every command… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset.texttext-generation1K<n<10K0 likes12 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.