datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.HumanCreativityBenchmark
The Human Creativity Benchmark (HCB)
Expert evaluations of AI-generated creative work, built to separate two signals that single-score benchmarks collapse: convergence, where professionals align around shared, checkable standards, and divergence, where creative taste legitimately differs. Each AI output is judged by domain professionals through three complementary lenses — forced-choice pairwise comparisons, 1-5 scalar ratings on prompt adherence, usability, and visual appeal… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/HumanCreativityBenchmark.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.human_curated_qa_dataset
Human Curated QA Dataset
DigiGreen/human_curated_qa_dataset is a human-verified question-answer dataset designed to support research and development in natural language question answering and agriculture-focused conversational AI.
This dataset contains realistic, domain-relevant QA pairs that were manually curated to ensure accurate and contextually rich answers. It can be used to benchmark models for QA generation.
📌 Dataset Overview
Name: Human Curated QA Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/human_curated_qa_dataset.The-android-and-the-human
Presentation
An English dataset composed of 10,000 real and 10,000 generated dreams.
Real dreams are from DreamBank. Generated dreams were produced with Oneirogen (0.5B, 1.5B and 7B), a language model for dream generation. Generation examples are available on my website.
This dataset can be used to determine differences between real and generated dreams. It can also be used to classify whether a dream narrative is generated or not.
This work was performed using HPC resources (Jean… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/The-android-and-the-human.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.humaneval_mt_pt
HumanEval-PT
Portuguese version of HumanEval, a code generation benchmark with programming problems and test cases.
Translated using NLLB-200.
Original Dataset: https://huggingface.co/datasets/openai/openai_humaneval
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/humaneval_mt_pt.Human-Like-Gut-Health-DPO-QnA
Gut Health DPO Dataset
Overview
This dataset contains 200 carefully curated examples for Direct Preference Optimization (DPO) training in the domain of gut health and digestive wellness. Each example consists of a user prompt, a "chosen" response (preferred), and a "rejected" response (less preferred), designed to train AI models to provide high-quality, medically responsible advice on digestive health topics.
Dataset Structure
The dataset is provided in CSV… See the full description on the dataset page: https://huggingface.co/datasets/Sid3503/Human-Like-Gut-Health-DPO-QnA.ro-human-machine-60kThe corpus for this study consists of multiple datasets of comparable text lengths, both machine-generated and human-written.
1401 books:
841 manually written abstracts provided by the Central University Library of Bucharest, representing descriptions of Romanian old documents (literary magazines and books dated between the 19th century and the present),
560 books descriptions (cartigratis.com, accessed 8 January 2024);
4320 news articles crawled from DigiNews (digi24.ro, accessed 8… See the full description on the dataset page: https://huggingface.co/datasets/readerbench/ro-human-machine-60k.
