text-detection
ai-text-detection-pile
Dataset Card for AI Text Dectection Pile
Dataset Summary
This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ.
Here is the (tentative) breakdown:
Human Text
Dataset
Num Samples
Link
Reddit WritingPromps
570k
Link
OpenAI Webtext
260k
Link
HC3 (Human Responses)
58k
Link
ivypanda-essays
TODO
TODO… See the full description on the dataset page: https://huggingface.co/datasets/artem9k/ai-text-detection-pile.handwritten_text_detection
Handwritten text detection dataset
Data domain
The blanks were provided by youth organization "Armenian Club" (telegram, instagram ), Russia Moscow.
The text on blanks was written during dictation "Teladrutyun" in 2018
The blanks were labeled by Amir and Renal during research project in HSE MIEM
Dataset info
Contains labeled dictations blanks in YOLO format
91 image in total, 73 (80%) for train and 18 (20%) for test
No image alignment or any preprocess… See the full description on the dataset page: https://huggingface.co/datasets/armvectores/handwritten_text_detection.ai-human-text-detection-v1
🧠 AI vs Human Text Detection Dataset (v1)
This dataset merges nine major public and academic corpora to form one of the most comprehensive resources for AI-generated text detection model training and evaluation.
🔗 Sources
The dataset consolidates, cleans, and standardizes multiple open datasets and research benchmarks, each focusing on human vs. AI-generated text classification:
Hello-SimpleAI / HC3 — Human–ChatGPT comparison corpus
gsingh1-py / train — Large-scale… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/ai-human-text-detection-v1.ocr-receipts-text-detectionThe Grocery Store Receipts Dataset is a collection of photos captured from various
**grocery store receipts**. This dataset is specifically designed for tasks related to
**Optical Character Recognition (OCR)** and is useful for retail.
Each image in the dataset is accompanied by bounding box annotations, indicating the
precise locations of specific text segments on the receipts. The text segments are
categorized into four classes: **item, store, date_time and total**.ai-text-detection-trainingNepali_Text_Detection_Dataset
