datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generated-csvshuman_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.AI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.Dynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.llm-generated-essaypersian-ai-generated-text
📝 Persian AI-Generated Text Dataset
A large-scale collection of 10,546 AI-generated Persian (Farsi) texts produced by 70 different large language models across diverse topics and writing styles. This dataset is designed to support research in AI-generated text detection for the Persian language.
Dataset Summary
Attribute
Value
Language
Persian (Farsi)
Total Samples
10,546
Unique Models
70
API Providers
9 (OpenRouter, NVIDIA, Free Endpoint, HF… See the full description on the dataset page: https://huggingface.co/datasets/ehsantorabi/persian-ai-generated-text.keystroke-dataset-raw-generatedgeneratedai-generated-questions-quality
🎓 IV WAPLA: AI Generated Questions Quality - What impacts quality and how to Improve Question Generation
The IV Workshop on Practical Applications of Learning Analytics and Artificial Intelligence in Brazil (Workshop de Aplicações Práticas de Learning Analytics em Instituições de Ensino no Brasil, WAPLA 2026) is a satellite event of the XV Brazilian Congress on Informatics in Education (Congresso Brasileiro de Informática na Educação, CBIE 2026).
In 2026 the 4th Edition of… See the full description on the dataset page: https://huggingface.co/datasets/aiboxlab/ai-generated-questions-quality.human_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both… See the full description on the dataset page: https://huggingface.co/datasets/jist2009/human_ai_generated_text.refinedweb-generated-questions
Generated Questions and Answers from the Falcon RefinedWeb Dataset
This dataset contains 1k open-domain questions and answers generated using documents from Falcon's refinedweb dataset using GPT-4. You can find more details about this work in the following blogpost.
Each row consits of:
document_id - an id of a text chunk from the refined web dataset, from which the question was generated. Each id contains the original document index from the refinedweb dataset, and the chunk index… See the full description on the dataset page: https://huggingface.co/datasets/pinecone/refinedweb-generated-questions.generated-critiques
Dataset Summary
This dataset presents critiques for 43k patches from 1,077 open-source python projects. The patches are generated by Nebius's agent and the critiques are generated by GPT-4o.
How to use:
from datasets import load_dataset
dataset = load_dataset("AGENTDARS/generated-critiques")
Software Patch Evaluation Dataset
This dataset contains structured information about software patches, problem statements, evaluations, and critiques. It is designed to assess… See the full description on the dataset page: https://huggingface.co/datasets/AGENTDARS/generated-critiques.ChatGPT-generated_fake_news_datasetgenerated-proofs-axioms
What machine-generated Lean proofs rest on
Per-theorem axiom dependencies for 9,729 machine-generated Lean 4 proofs.
A model writing Lean gets one bit of feedback: the proof compiles, or it does
not. What the proof ends up standing on is not part of that signal. This is that
measurement, over the Goedel-Prover output for the Lean Workbook problems.
Headline
Of the 9,169 proofs that still compile under Lean 4.32:
count
share
reach Classical.choice
8,496… See the full description on the dataset page: https://huggingface.co/datasets/vince-gonzalez/generated-proofs-axioms.agent_questions_generated
Agent Questions Generated Pt-Br
Detalhes do Dataset
Descrição
Desenvolvidor por:
Álvaro Lopes. Linkedin
Artur de Vlieger Linkedin
Fabrício Salomon Linkedin
Leticia Bossatto Marchezi Linkedin
Luis Felipe Jorge Linkedin
Otávio Coletti Linkedin
Patrocinado por : Pico
Língua(s) (NLP) :Português
Correspondência: raia.projetos@gmail.com, leticiabossatto@gmail.com
Fontes
Repositório: Github
Uso
O dataset pode ser utilizado… See the full description on the dataset page: https://huggingface.co/datasets/letmarchezi/agent_questions_generated.Generated_Restaurant_Reviews_GPT3.5
license: cc-by-4.0
task_categories:
text-classification
language:
tr
tags:
food
Generated Review
size_categories:
1K<n<10K
augmented_dataset_llm_generated_NER
📚 Augmented LLM-Generated NER Dataset for Scholarly Text
🧠 Dataset Summary
This dataset contains synthetically generated academic text tailored for Named Entity Recognition (NER) in the software engineering domain. The synthetic data augments scholarly writing using large language models (LLMs), with entity consistency maintained via token preservation.
The dataset is generated by merging and rephrasing pairs of annotated sentences from scholarly papers using… See the full description on the dataset page: https://huggingface.co/datasets/psresearch/augmented_dataset_llm_generated_NER.generated-spans-detection
Dataset description
This dataset is designed for training and evaluating models tasked with detecting text fragments generated by Large Language Models (LLMs) within written scientific discourse.
Data generation
Abstracts of scientific articles from journals indexed in Higher Attestation Commission list were used as the source material for generating synthetic examples. Abstracts were chosen due to their high information density and specialized terminology, allowing for… See the full description on the dataset page: https://huggingface.co/datasets/iis-research-team/generated-spans-detection.pepforge-generated-data
PepForge — Generated Peptide Library
Large-scale generated peptide library from PepForge's hierarchical cascade pipeline (Layout GPT → Content GPT-L → Connection GAT-L), with AMP activity prediction and ADMET profiling.
Dataset Summary
Metric
Value
Total novel molecules
4,783,266
Generation
10M raw samples (5 shards × 2M)
Deduplication
InChIKey-based: removed exact duplicates + 246,734 training-set overlaps (training corpus = 383,817 molecules)… See the full description on the dataset page: https://huggingface.co/datasets/qingxin1999/pepforge-generated-data.mobileforge-generated-tasks
MobileForge Generated Tasks
This dataset contains the consolidated task pool generated by MobileGym-Curriculum from target-app exploration trajectories. These tasks are used by MobileForge for rollout collection and annotation-free adaptation.
Dataset summary
File
Rows
Apps
Size
Description
generated_tasks_26020301-all.csv
3,249
20
1.93 MB
Consolidated AndroidWorld-side MobileForge task pool.
The task pool is generated from real target-app… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-generated-tasks.ai-vs-real-text-generated-datasetGenerated_News_Political_LeaningSynthetic_Story_Generated_By_Geminimanagpt-4080-nlp-prompts-and-generated-textsThis dataset includes 4,080 texts that were generated by the ManaGPT-1020 large language model, in response to particular input sequences.
ManaGPT-1020 is a free, open-source model available for download and use via Hugging Face’s “transformers” Python package. The model is a 1.5-billion-parameter LLM that’s capable of generating text in order to complete a sentence whose first words have been provided via a user-supplied input sequence. The model represents an elaboration of GPT-2 that has… See the full description on the dataset page: https://huggingface.co/datasets/NeuraXenetica/managpt-4080-nlp-prompts-and-generated-texts.M1_MCQ_generated_with_questionsenriched-generated-arguments
Info
This is a version of a generated arguments corpus enriched with linguistic features and argument quality dimensions.
The linguistic features were extracted with elfen.
The argument quality dimensions were extracte with these adapters.
Citation
If you use this enriched version of the generated arguments corpus, please cite
@inproceedings{doenmez-maurer-2025-ai,
title = "AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive… See the full description on the dataset page: https://huggingface.co/datasets/mmmaurer/enriched-generated-arguments.pdv_generated_qa_final_users_1generated_lecturesgenerated
