datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-triplet-lifestyle-romance
🏯 Japanese Preference Dataset: Counseling & Advice (Free Sample)
This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment.
The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.… See the full description on the dataset page: https://huggingface.co/datasets/wasabiP/japanese-triplet-lifestyle-romance.FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated
Federico García Lorca - Annotated Poetry Dataset
A curated and annotated dataset of 283 poems by Federico García Lorca, spanning 9 of his major works (1921--1940). Each poem is enriched with publication metadata and GPT-4-generated thematic and contextual annotations.
Use Case: LLM Generalization Evaluation
This dataset was created to evaluate how well large language models can generalize literary style from a small, domain-specific corpus. It has been used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/xaviviro/FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated.gpt5.6_romance_multi_turn
GPT-5.6 Romance Multi-Turn
Private English romance roleplay training data generated by GPT-5.6 and independently validated and rechecked before approval.
Contents
conversations.jsonl: 4,000 approved assistant-prefix targets.
manifest.json: source, expansion, review thresholds, row counts, and SHA-256 checksums.
500 source conversations expanded into 8 targets each.
47 concise reasoning targets and 3,953 direct-prose targets.
Data contract
The… See the full description on the dataset page: https://huggingface.co/datasets/Skttttt/gpt5.6_romance_multi_turn.
