CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01diversen /leaves-of-grass leaves of grass The following is a dataset for training a model to generate text in the style of Walt Whitman's "Leaves of Grass". The idea with this dataset is to provide a single line (input) and then provide the next lines (1 to 10 lines) of the poem as output. A model can then be trained to generate lines of poems given a single line of input. There is a generate_data.py script that can be used to generate the dataset. It keeps some formatting. New lines may be indented by a… See the full description on the dataset page: https://huggingface.co/datasets/diversen/leaves-of-grass.textquestion-answering10K<n<100K0 likes317 downloads2y agoHugging Face02gretelai /gsm8k-synthetic-diverse-8b gretelai/gsm8k-synthetic-diverse-8b This dataset is a synthetically generated version inspired by the GSM8K https://huggingface.co/datasets/openai/gsm8k dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-8B as the agent LLM. It contains ~1500 Grade School-level math word problems with step-by-step solutions, focusing on age group, difficulty, and domain diversity. Key Features: Synthetically Generated: Math problems created using Gretel… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gsm8k-synthetic-diverse-8b.textquestion-answering1K<n<10K0 likes300 downloads2y agoHugging Face03Naholav /CodeGen-Diverse-5K CodeGen-Diverse-5K: Broad Coverage for Competitive Programming Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset) Dataset Description CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions. Key Statistics Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.tabulartext-generation1K<n<10K0 likes100 downloads10mo agoHugging Face04kunu5402 /Diverse-Knowledge Everything Data This data is synthetically generated by a ton of open and closed source models. This is basically a parsed version of yearly log form a small dialouge based testing to anylyze model's response on it then perform human evals on it. The data contains information about everything from every domain, most of the pairs included in this data are preferred by humans as the model's response. It can be used for topic modeling, or human preference evals etc. Rest anyone can do… See the full description on the dataset page: https://huggingface.co/datasets/kunu5402/Diverse-Knowledge.texttext-generation10K<n<100K0 likes41 downloads2y agoHugging Face05datasetter458 /hollow-knight-diverse-dataset Hollow knight diverse dataset A dataset containing every tiny detail about the game 'hollow knight'. date of the wiki data : 5/31/2026 columns : "Question", "Answer" composition : comprised of Q&A pairs for every page of the hollow knight wiki, all the 557 pages. reference :all_pages textquestion-answering1K<n<10K0 likes27 downloads4mo agoHugging Face06ceselder /loracle-ia-diverse-qa loracle-ia-diverse-qa — v6 QA training data for the loracle — a model that reads a LoRA's weight deltas and answers questions about the behavior it encodes. What this is Each row pairs a LoRA identifier with a (question, answer) where the answer requires reading the LoRA's direction-token projections to answer correctly. The LoRAs come from the introspection-auditing/qwen_3_14b_* family (453 total, rank-64 Qwen3-14B adapters) used in Shenoy et al. (2026) Introspection… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-ia-diverse-qa.textquestion-answering1K<n<10K0 likes13 downloads6mo agoHugging Face07datasetter458 /diverse-bash-dataset-extended Bash diverse dataset extended This dataset is a continious work from datasetter458/bash-diverse-dataset Commands covered: bash, apt, crontab, date, diff, docker, env, export, fdisk, file, fsck, gcc, git, history, join, kubectl, lsof, man, makefs, mount, nice, node, nohup, npm, passwd, paste, patch, pkill, python, renice, split, ss, time, umount, uname, uniq, vi, whereis, which Note : For each command, a big amount of Q&As, and each set of them covers pretty much all the flags for… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset-extended.texttext-generation1K<n<10K0 likes13 downloads4mo agoHugging Face08nattiey1 /diverse-unit-QA Dataset Card for DUQA Dataset Description Abstract DUQA is a dataset for single-step unit conversion questions. It comes in three sizes, ”DUQA10k”, ”DUQA100k” and ”DUQA1M”, with 10,000, 100,000 and 1,000,000 entries respectively. Each size contains a mixture of basic and complex conversion questions, including simple conversion, multiple answer, max/min, argmax/argmin, and noisy/q-noisy questions. The complexity level varies based on the amount of information… See the full description on the dataset page: https://huggingface.co/datasets/nattiey1/diverse-unit-QA.textquestion-answering100K<n<1M0 likes10 downloads3y agoHugging Face09datasetter458 /diverse-bash-dataset Diverse bash terminal dataset A dataset containing Question&Code pairs for most of the standard bash commands. columns : "Question", "Code answer". Commands covered : echo, cat, cd, rm, mkdir, top, free, du, df, ps, head, tail, grep, cp, cut, sort, touch, ls, groupadd, ifconfig, ip, ln, ping, scp, ssh, sudo, systemctl, tar, useradd, userdel, usermod, wc specs : high quality dataset with precise coding examples for each command, covering (almost) all the flags of every command… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset.texttext-generation1K<n<10K0 likes9 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.