datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.mobileforge-generated-tasks
MobileForge Generated Tasks
This dataset contains the consolidated task pool generated by MobileGym-Curriculum from target-app exploration trajectories. These tasks are used by MobileForge for rollout collection and annotation-free adaptation.
Dataset summary
File
Rows
Apps
Size
Description
generated_tasks_26020301-all.csv
3,249
20
1.93 MB
Consolidated AndroidWorld-side MobileForge task pool.
The task pool is generated from real target-app… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-generated-tasks.controlled-generated-convos-gpt-4.1-mini
Controlled Generated Conversations: gpt-4.1-mini
Dataset Description
This dataset contains synthetic customer support conversations generated using gpt-4.1-mini as part of research on cross-lingual stability of LLM judges. The conversations are designed for evaluating how well language models maintain consistent performance across different languages, with a focus on Finno-Ugric languages (Estonian, Finnish, Hungarian) and English.
Dataset Summary
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/isaacchung/controlled-generated-convos-gpt-4.1-mini.pepforge-generated-data
PepForge — Generated Peptide Library
Large-scale generated peptide library from PepForge's hierarchical cascade pipeline (Layout GPT → Content GPT-L → Connection GAT-L), with AMP activity prediction and ADMET profiling.
Dataset Summary
Metric
Value
Total novel molecules
4,783,266
Generation
10M raw samples (5 shards × 2M)
Deduplication
InChIKey-based: removed exact duplicates + 246,734 training-set overlaps (training corpus = 383,817 molecules)… See the full description on the dataset page: https://huggingface.co/datasets/qingxin1999/pepforge-generated-data.enriched-generated-arguments
Info
This is a version of a generated arguments corpus enriched with linguistic features and argument quality dimensions.
The linguistic features were extracted with elfen.
The argument quality dimensions were extracte with these adapters.
Citation
If you use this enriched version of the generated arguments corpus, please cite
@inproceedings{doenmez-maurer-2025-ai,
title = "AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive… See the full description on the dataset page: https://huggingface.co/datasets/mmmaurer/enriched-generated-arguments.
