datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds.
We generate the dataset with the following steps and two approaches:
Generate ~110k descriptions by GPT4o.
Approach 1: Generate ~110k codes follow each description by GPT4o-mini.
Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions.
Run the ~220k codes and do auto-filtering.
Get the final ~200k legitimate ARC-like tasks with examples.
task1729_personachat_generate_next
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1729_personachat_generate_next
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1729_personachat_generate_next.ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4o-mini.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
generate-narrate-tinystories-pretrain
Narrative · TinyStories · Pretraining (Cleaned)
Microsoft's TinyStories V2, cleaned and stored as parquet. 2,745,100 stories, 441 million
words, one story per row with provenance on every record.
Composition
Config
Records
%
Source
all
2,745,100
100.00
the single config (default)
gpt-4
2,745,100
100.00
TinyStoriesV2-GPT4-train
TinyStories V2 holds samples generated by GPT-3.5 and samples generated by GPT-4. Only
the GPT-4 samples are here… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/generate-narrate-tinystories-pretrain.qwen-generated-svamp-controls-sft
Qwen-Generated SVAMP CoT Controls ? SFT
Qwen-generated controlled reasoning traces for SVAMP in LLaMA-Factory SFT format. Variants include ordinary, all-caps, no-comma, disclaimer, and multilingual examples.
Splits
3,940 training examples and 380 held-out evaluation examples.
Format
The JSON files use the LLaMA-Factory Alpaca-style schema. The included
dataset_info.json registers the exact training and evaluation names. DPO
records are marked with… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-svamp-controls-sft.task957_e2e_nlg_text_generation_generate
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task957_e2e_nlg_text_generation_generate
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task957_e2e_nlg_text_generation_generate.warp-taskgen-generated-ipi-tasks-50
WARP Taskgen Generated IPI Tasks 50
Dataset Summary
This dataset contains WARP Taskgen Phase 4 browser-agent trajectories for a
50-task generated indirect prompt injection (IPI) cohort. The trajectories were
produced with the AgentLab harness on
WebArena GitLab and Postmill (Reddit) benchmark applications.
The export is a report-only projection of already written benchmark artifacts.
It does not alter scoring, PVPO encounter checks, rewards, admission, or
trajectory… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset-submission-warp/warp-taskgen-generated-ipi-tasks-50.Generated_OE_Gregory_Dialogues_Text_and_Evaluation
Generated Old English Gregory's Dialogues (variatio)
A complete, machine-generated Old English variatio of the Old English Dialogues of
Gregory the Great (Waerferth's translation), produced on 19 July 2026, together with the
full generation and evaluation apparatus: prompt, constraint lexicon scripts, validator,
dependency parses, word embeddings, and all quantitative evaluation results.
The project is described in:
Martin Arista, J., & Nunez, M. Evaluating Generated Old… See the full description on the dataset page: https://huggingface.co/datasets/Nerthus-Project/Generated_OE_Gregory_Dialogues_Text_and_Evaluation.Dynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.task389_torque_generate_temporal_question
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task389_torque_generate_temporal_question
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task389_torque_generate_temporal_question.ghanaian-corpus-generate-clean
Ghanaian Corpus Generate — Cleaned
Cleaned version of ghananlpcommunity/ghanaian-corpus-generate.
What changed
The text_clean column was produced by stripping out non-sentence content from text:
Removed section/chapter numbering (e.g. 24 5.3, III, 6.1)
Removed document headers (e.g. Chapter 6: Conclusion and Recommendations, Abstract, Acknowledgements)
Removed figure/table/code references (e.g. Figure 4.10 QR recognition class)
Removed very short fragments (< 15… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghanaian-corpus-generate-clean.generated-e-mail-spamThe dataset consists of a **CSV file** containing of 300 generated email spam messages.
Each row in the file represents a separate email message, its *title and text.*
The dataset aims to facilitate the analysis and detection of spam emails.
The dataset can be used for various purposes, such as *training machine learning
algorithms to classify and filter spam emails, studying spam email patterns,
or analyzing text-based features of spam messages*.Generated-Empathetic-Dialogues-v0.1-Smol
Generated Empathetic Conversations v0.1 - Smol
This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics.
Highlights
Multi-round conversation
It's not single-turn. The user and the assistant works together to gradually unfold the conversation.
The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.ai-generated-texts
Spanish DPO Preference Pairs for Detector Evasion
Preference pairs for DPO fine-tuning of Qwen/Qwen2.5-0.5B-Instruct against the Oculus multilingual AI text detector on Spanish academic abstracts. Repository id: pymlex/ai-generated-texts.
Dataset size
Statistic
Count
Train abstracts processed
8891
DPO pairs retained
6396
Pairs skipped by logit margin
2495
Empty paraphrase pairs
0
Logit margin threshold: absolute gap at least 1.… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/ai-generated-texts.mobileforge-generated-tasks
MobileForge Generated Tasks
Anonymous project: https://mobileforge-anonymous.github.io/Anonymous code: https://github.com/mobileforge-anonymous/MobileForge
This dataset contains the consolidated task pool generated by MobileGym-Curriculum from target-app exploration trajectories. These tasks are used by MobileForge for rollout collection and annotation-free adaptation.
Release inventory payloads: files=1; bytes=2028399;… See the full description on the dataset page: https://huggingface.co/datasets/mobileforge-anonymous/mobileforge-generated-tasks.qwen-generated-svamp-controls-dpo
Qwen-Generated SVAMP CoT Controls ? DPO
Preference pairs built from Qwen-generated SVAMP reasoning traces in LLaMA-Factory DPO format. Each record contains instruction, input, chosen, and rejected fields.
Splits
3,152 training preference pairs and 304 held-out evaluation pairs.
Format
The JSON files use the LLaMA-Factory Alpaca-style schema. The included
dataset_info.json registers the exact training and evaluation names. DPO
records are marked with… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-svamp-controls-dpo.ruleloopvit-sft-generated-rules-019200
RuleLoopViT SFT Generated ARC-AGI-1 Rules
This dataset contains one generated rule text for each of the 400 ARC-AGI-1
training tasks.
The rules were generated by the RuleLoopViT principal SFT language model:
omrisap/sft_lm_principal
Each task was prompted with 2-4 official ARC-AGI demonstration pairs and decoded
deterministically. The generated output follows the project's five-section rule
schema:
[CORE RULE]
[INPUT STRUCTURE]
[TARGET SELECTION]
[TRANSFORMATION]
[OUTPUT… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/ruleloopvit-sft-generated-rules-019200.full-html-stying-dataset-generated-css-from-style-plan
Generated CSS From Style Plan
kogai/full-html-stying-dataset-generated-css-from-style-plan contains generated_css_from_style_plan.jsonl, a JSONL dataset with 44458 synthetic examples. Model-generated CSS outputs conditioned on source HTML, user style requests, and structured style plans.
Schema
chat_template_overhead_tokens: field present in the JSONL records.
created_at: field present in the JSONL records.
input_html: source HTML before Tailwind classes are… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-generated-css-from-style-plan.diffusion-generated-text
Diffusion-Generated Text
This dataset contains 16,791 question-response pairs generated by a diffusion language model. It is released to support research on diffusion-generated language and machine-generated text detection.
Dataset schema
Column
Type
Description
question
string
Input question or prompt.
dLLM_response
string
Response produced by the diffusion language model.
Loading
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/paoche11/diffusion-generated-text.entity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-sft-dataset-GPT-4.0-generated-v1.entity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment… See the full description on the dataset page: https://huggingface.co/datasets/fibonacciai/entity-attribute-sft-dataset-GPT-4.0-generated-v1.mobileforge-generated-tasks
MobileForge Generated Tasks
This dataset contains the consolidated task pool generated by MobileGym-Curriculum from target-app exploration trajectories. These tasks are used by MobileForge for rollout collection and annotation-free adaptation.
Dataset summary
File
Rows
Apps
Size
Description
generated_tasks_26020301-all.csv
3,249
20
1.93 MB
Consolidated AndroidWorld-side MobileForge task pool.
The task pool is generated from real target-app… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-generated-tasks.ultrachat_200k_generated_gemma-2-2b-itThis dataset contains 512 answers generated by the gemma-2-2b-it model on a subset of the ultrachat 200k test_sft dataset using greedy decoding.
The subset was generated by filtering out conversations that were >= 1024 - 128 tokens long, and answers were cut off at each batch after 1024 - min(batch_prompt_lengths) generated tokens, such that each answer is at most 128 tokens long. The generated answers are 200k tokens so 390 tokens (~300 words or 2/3 pages) on average.
Ai_generate_1controlled-generated-convos-gpt-4.1-mini
Controlled Generated Conversations: gpt-4.1-mini
Dataset Description
This dataset contains synthetic customer support conversations generated using gpt-4.1-mini as part of research on cross-lingual stability of LLM judges. The conversations are designed for evaluating how well language models maintain consistent performance across different languages, with a focus on Finno-Ugric languages (Estonian, Finnish, Hungarian) and English.
Dataset Summary
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/isaacchung/controlled-generated-convos-gpt-4.1-mini.qwen-generated-h4-controls-5k-sft
Qwen-Generated H4 CoT Controls ? 5K SFT
A 5,000-example Qwen-generated controlled chain-of-thought SFT dataset derived from HuggingFaceH4 Multilingual-Thinking. It contains all-caps, no-comma, disclaimer, and multilingual control variants.
Splits
4,500 training examples and 500 held-out evaluation examples.
Format
The JSON files use the LLaMA-Factory Alpaca-style schema. The included
dataset_info.json registers the exact training and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-h4-controls-5k-sft.entity-attribute-dataset-GPT-3.5-generated-v1
Entity Attribute Dataset 306k (GPT-3.5 generated)
Dataset Summary
The Entity Attribute Dataset 306k (GPT-3.5 generated) is designed for instruction fine-tuning, specifically for the task of generating structured catalogs in JSON format based on product titles. The dataset includes a diverse range of products from various categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and more.
Usage
This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-dataset-GPT-3.5-generated-v1.task036_qasc_topic_word_to_generate_related_fact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task036_qasc_topic_word_to_generate_related_fact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task036_qasc_topic_word_to_generate_related_fact.
