datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RHyMENeXTVideoA copy of the NextQA dataset, used to demonstrate fine-tuning Aria on video dataset.
magenta-realtime-mlx-cpp
Magenta RealTime — C++ MLX runtime bundle
This dataset is a re-packaging of
Google's Magenta RealTime weights
for the C++ MLX runtime in
rhymeswithlion/magenta-realtime-mlx-cpp.
It contains exactly what mlx-stream needs at startup; nothing more, nothing
less. The upstream .pt / .npy checkpoints are intentionally not
mirrored here — they're only useful for the (Python) re-export tooling on the
project's main distribution.
Contents
.
├──… See the full description on the dataset page: https://huggingface.co/datasets/rhymeswithlion/magenta-realtime-mlx-cpp.NLVR2A copy of the NLVR2 dataset, used to demonstrate fine-tuning Aria on multi-image dataset.
task183_rhyme_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task183_rhyme_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task183_rhyme_generation.RefCOCOA subset of the RefCOCO dataset, used to demonstrate fine-tuning Aria.
gemma4-rhyme-interp
Gemma-4 Rhyme Interpretability — datasets
Evaluation and probing datasets from a mechanistic interpretability study of how
google/gemma-4-E2B (base) completes the last word of a rhyming line. The full
analysis, code, and write-ups (reports 01–10, including the circuit, the
localization of the rhyme "write" to a single MLP, the value-memory readout, and
a training-free rank-1 weight edit that installs a false rhyme) live in the
GitHub repository:… See the full description on the dataset page: https://huggingface.co/datasets/eac123/gemma4-rhyme-interp.rhyme-sentencesrhymes-ai__Aria-details
Dataset Card for Evaluation run of rhymes-ai/Aria
Dataset automatically created during the evaluation run of model rhymes-ai/Aria
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rhymes-ai__Aria-details.Formal_Poetry_Rhyme_PairsA re-curated selection of English-language poetry rhyme pairs used in the dataset/article linked below (and initially drawn from the Chadwyck-Healey corpus).
Link to the full source data
SOURCE:
Dataset for "Generative Aesthetics: On the formal stuckness of AI verse"
(Journal of Cultural Analytics_, vol. 10, no. 3, Sept. 2025)
https://culturalanalytics.org/article/id/1036/
https://doi.org/10.7910/DVN/BEQAYG
nursery_rhyme_fgo
Dataset of nursery_rhyme/ナーサリー・ライム/童谣 (Fate/Grand Order)
This is the dataset of nursery_rhyme/ナーサリー・ライム/童谣 (Fate/Grand Order), containing 401 images and their tags.
The core tags of this character are long_hair, bow, hat, white_hair, black_headwear, purple_eyes, black_bow, beret, very_long_hair, striped_bow, braid, hair_bow, twin_braids, pink_eyes, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/nursery_rhyme_fgo.flan_combined_task183_rhyme_generationShevchenko_rhymed_RU_UKRrhymer_toneshtest-rhymemy_first_lora_v1-datasetalicealice2test_words_aozora_rhymepoetry-rhymehuman-performance-2026
