datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaves-of-grass
leaves of grass
The following is a dataset for training a model to generate text in the style of Walt Whitman's "Leaves of Grass".
The idea with this dataset is to provide a single line (input) and then provide the next lines (1 to 10 lines) of the poem as output.
A model can then be trained to generate lines of poems given a single line of input.
There is a generate_data.py script that can be used to generate the dataset.
It keeps some formatting. New lines may be indented by a… See the full description on the dataset page: https://huggingface.co/datasets/diversen/leaves-of-grass.CodeGen-Diverse-5K
CodeGen-Diverse-5K: Broad Coverage for Competitive Programming
Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset)
Dataset Description
CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions.
Key Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.diverse-svg-prompts
Diverse SVG Prompts
Diverse SVG Prompts is a public collection of 20,000 high-quality,
generated and filtered English briefs for SVG and vector-graphics generation.
It contains 18,000 general illustration prompts and 2,000 lettering prompts.
Schema
The dataset intentionally has only two columns:
prompt: the complete visual brief.
type_tags: a list of category, author-model, and processing tags.
Example:
{
"prompt": "A moonlit mechanical heron..."… See the full description on the dataset page: https://huggingface.co/datasets/Nbardy/diverse-svg-prompts.dualmsm-cheese-mixes-diverse
dualmsm-cheese-mixes-diverse
Two finetune-ready cheese-preference mixtures for the dual-MSM cheese dissociation experiments, freshly assembled from the diverse cheese-AFT datasets (the original small sets plus the expanded sets). Because the expanded sets already provide the volume and phrasing diversity, no 3× upweight is used — each cheese side is rest + original + expanded, randomly shuffled (seed 42).
file
rows
teaches
rest_amercheese_diverse.jsonl
29,899
like… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-mixes-diverse.BERnaT-Diverse
BERnaT: Basque Encoders for Representing Natural Textual Diversity
Submitted to LREC 2026
Abstract
Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally
exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this
paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal,
historical, informal, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BERnaT-Diverse.
