datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-text-detection-pile-cleaned
AI Text Detection Pile - Cleaned Dataset
Dataset Description
This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training.
Dataset Details
Total Samples: 721,626 (cleaned from… See the full description on the dataset page: https://huggingface.co/datasets/R-obi/ai-text-detection-pile-cleaned.ai_text_detection_dataset_dl_hw_2_v6
