datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.AI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.human-vs-Ai-generated-datasetai-generated-ecommerce-damaged-productai-generated-ecommerce-fake-logisticsai-generated_videoai-generated-ecommerce-damaged-electronicsai-generated-songsjapanese-AI-generatednbeerbower/japanese-photos-captioned captions used to generate images with Tongyi-MAI/Z-Image-Turbo
ai-generated-songs2ai-generated-texts
Spanish DPO Preference Pairs for Detector Evasion
Preference pairs for DPO fine-tuning of Qwen/Qwen2.5-0.5B-Instruct against the Oculus multilingual AI text detector on Spanish academic abstracts. Repository id: pymlex/ai-generated-texts.
Dataset size
Statistic
Count
Train abstracts processed
8891
DPO pairs retained
6396
Pairs skipped by logit margin
2495
Empty paraphrase pairs
0
Logit margin threshold: absolute gap at least 1.… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/ai-generated-texts.ai-generated-text-classification
Dataset Card for "ai-generated-text-classification"
More Information needed
iceland-AI-generatednbeerbower/iceland-photos captions used to generate images with Tongyi-MAI/Z-Image-Turbo
human-ai-generated-text
Dataset Card for human-ai-generated-text
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/ardavey/human-ai-generated-text/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ardavey/human-ai-generated-text.human_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both… See the full description on the dataset page: https://huggingface.co/datasets/jist2009/human_ai_generated_text.ai-generated-questions-quality
🎓 IV WAPLA: AI Generated Questions Quality - What impacts quality and how to Improve Question Generation
The IV Workshop on Practical Applications of Learning Analytics and Artificial Intelligence in Brazil (Workshop de Aplicações Práticas de Learning Analytics em Instituições de Ensino no Brasil, WAPLA 2026) is a satellite event of the XV Brazilian Congress on Informatics in Education (Congresso Brasileiro de Informática na Educação, CBIE 2026).
In 2026 the 4th Edition of… See the full description on the dataset page: https://huggingface.co/datasets/aiboxlab/ai-generated-questions-quality.ai-generated-ecommerce-damaged-phone-screenai-generated-ecommerce-fragile-brokenai-generated-songs3persian-ai-generated-text
📝 Persian AI-Generated Text Dataset
A large-scale collection of 10,546 AI-generated Persian (Farsi) texts produced by 70 different large language models across diverse topics and writing styles. This dataset is designed to support research in AI-generated text detection for the Persian language.
Dataset Summary
Attribute
Value
Language
Persian (Farsi)
Total Samples
10,546
Unique Models
70
API Providers
9 (OpenRouter, NVIDIA, Free Endpoint, HF… See the full description on the dataset page: https://huggingface.co/datasets/ehsantorabi/persian-ai-generated-text.generated-ai-sample
Dataset Card for "generated-ai-sample"
More Information needed
toxic_AI_generatedai-vs-real-text-generated-datasetAI-Generated_Chinese_Modern_Poetry
中文
这个数据集是使用DeepSeek-R1根据标题或摘要生成中文现代诗而构成的。
我们使用这个数据集训练了芝麻Zhima。芝麻是一个专注于中文现代诗创作的LLM,能根据用户指令用标题、摘要或关键词生成原创中文现代诗。
芝麻的名字来源于志摩(徐志摩)的谐音。徐志摩(1897-1931)是一位著名的中国现代诗人。
致谢
感谢modern-poetry,我们使用该项目中汇总的中文现代诗的标题和提炼出的摘要。
English
This dataset consists of Chinese modern poems generated by DeepSeek-R1 based on titles or summaries.
We use this dataset to train Zhima. Zhima is an LLM focused on Chinese modern poetry creation, capable of generating original Chinese modern poems based… See the full description on the dataset page: https://huggingface.co/datasets/Hyaline/AI-Generated_Chinese_Modern_Poetry.AI_Human_generated_movie_reviews
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The "AI_Human_generated_movie_reviews" dataset consists of 5.23k AI-generated movie reviews alongside 5.23k human-written reviews from the Stanford IMDB dataset. The AI reviews were created using several models, including Gemini 1.5 Pro, GPT-3.5-Turbo, and GPT-4.0-Turbo-Preview, via the… See the full description on the dataset page: https://huggingface.co/datasets/Lyra-stellAI/AI_Human_generated_movie_reviews.tupler-ai-generatedai-generated-chat-dataset
AI-Generated Chat Dataset
This public dataset contains 928 short user/assistant dialogue examples converted from dataset.md.
Provenance
The user questions/prompts were sourced from VMware/open-instruct. The assistant responses were AI-generated with google/gemma-4-12B.
Important Notice
This dataset is AI-generated. It may contain unintended wording, inaccuracies, biases, sensitive topics, or phrasing that does not reflect anyone's values or… See the full description on the dataset page: https://huggingface.co/datasets/Abhiram1009/ai-generated-chat-dataset.gemma4-e2b-generated-instructions-demo-v1
Unsloth Dataset Workflow Test
Overview
This dataset is a workflow validation dataset generated using Unsloth Studio.
It demonstrates the complete pipeline:
Source dataset
AI-generated instructions
Export to Parquet
Upload to Hugging Face
Dataset viewer validation
This repository is intended for testing the publication workflow before creating a larger production-quality dataset.
Dataset Structure
Columns
output
generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.tupler-ai-generatedgenerated_data_alpaca
