CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artem9k /ai-text-detection-pile Dataset Card for AI Text Dectection Pile Dataset Summary This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ. Here is the (tentative) breakdown: Human Text Dataset Num Samples Link Reddit WritingPromps 570k Link OpenAI Webtext 260k Link HC3 (Human Responses) 58k Link ivypanda-essays TODO TODO… See the full description on the dataset page: https://huggingface.co/datasets/artem9k/ai-text-detection-pile.text1M<n<10M46 likes591 downloads4y agoHugging Face02armvectores /handwritten_text_detection Handwritten text detection dataset Data domain The blanks were provided by youth organization "Armenian Club" (telegram, instagram ), Russia Moscow. The text on blanks was written during dictation "Teladrutyun" in 2018 The blanks were labeled by Amir and Renal during research project in HSE MIEM Dataset info Contains labeled dictations blanks in YOLO format 91 image in total, 73 (80%) for train and 18 (20%) for test No image alignment or any preprocess… See the full description on the dataset page: https://huggingface.co/datasets/armvectores/handwritten_text_detection.imageobject-detection10K<n<100K7 likes417 downloads2y agoHugging Face03silentone0725 /ai-human-text-detection-v1 🧠 AI vs Human Text Detection Dataset (v1) This dataset merges nine major public and academic corpora to form one of the most comprehensive resources for AI-generated text detection model training and evaluation. 🔗 Sources The dataset consolidates, cleans, and standardizes multiple open datasets and research benchmarks, each focusing on human vs. AI-generated text classification: Hello-SimpleAI / HC3 — Human–ChatGPT comparison corpus gsingh1-py / train — Large-scale… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/ai-human-text-detection-v1.text10K<n<100K8 likes191 downloads11mo agoHugging Face04UniqueData /ocr-receipts-text-detectionThe Grocery Store Receipts Dataset is a collection of photos captured from various **grocery store receipts**. This dataset is specifically designed for tasks related to **Optical Character Recognition (OCR)** and is useful for retail. Each image in the dataset is accompanied by bounding box annotations, indicating the precise locations of specific text segments on the receipts. The text segments are categorized into four classes: **item, store, date_time and total**.textimage-to-text14 likes165 downloads1y agoHugging Face05coai /ai-text-detection-trainingtext10K<n<100K3 likes158 downloads9mo agoHugging Face06Rajan /Nepali_Text_Detection_Datasetimage10K<n<100K1 likes150 downloads2y agoHugging Face07DonkeySmall /Yolo-Text-DetectionA text scene detection dataset, manually labeled in yolo format and contains 25 197 images, mostly with Latin and Russian text. image-to-text10K<n<100K4 likes150 downloads9mo agoHugging Face08UniqueData /ocr-text-detection-in-the-documents OCR Text Detection in the Documents Object Detection dataset The dataset is a collection of images that have been annotated with the location of text in the document. The dataset is specifically curated for text detection and recognition tasks in documents such as scanned papers, forms, invoices, and handwritten notes. The dataset contains a variety of document types, including different layouts, font sizes, and styles. The images come from diverse sources, ensuring a representative… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/ocr-text-detection-in-the-documents.imageimage-to-textn<1K9 likes129 downloads1y agoHugging Face09srikanthgali /ai-text-detection-pile-cleaned AI Text Detection Pile - Cleaned Dataset Dataset Description This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training. Dataset Details Total Samples: 721,626 (cleaned from original… See the full description on the dataset page: https://huggingface.co/datasets/srikanthgali/ai-text-detection-pile-cleaned.texttext-classification100K<n<1M5 likes113 downloads1y agoHugging Face10prosa-text /climate-stance-detectiontext10K<n<100K2 likes110 downloads2y agoHugging Face11amaye15 /receipts-text-detectionimage10K<n<100K3 likes99 downloads2y agoHugging Face12Mharis205 /ai-text-detection-pile Dataset Card for AI Text Dectection Pile Dataset Summary This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ. Here is the (tentative) breakdown: Human Text Dataset Num Samples Link Reddit WritingPromps 570k Link OpenAI Webtext 260k Link HC3 (Human Responses) 58k Link ivypanda-essays TODO TODO… See the full description on the dataset page: https://huggingface.co/datasets/Mharis205/ai-text-detection-pile.text1M<n<10M0 likes90 downloads9mo agoHugging Face13Roxanne-WANG /AI-Text_Detection About This Dataset This dataset is derived from the dataset used in the SeqXGPT paper and is uploaded solely for academic and research purposes related to our course project. The dataset is provided as-is and should not be used for any commercial applications or unauthorized activities. Our objective is to facilitate reproducibility and further research in AI-generated text (AIGT) detection. Here we upload part of its raw data and processed data by using gen_features.py in SeqXGPT… See the full description on the dataset page: https://huggingface.co/datasets/Roxanne-WANG/AI-Text_Detection.text-classification1K<n<10K1 likes79 downloads2y agoHugging Face14petersunde /manga-covers-text-detection Manga Covers Text Detection Manually annotated text boxes for manga covers text detection created by @JustANormalTinkerer. This dataset contains 100 cover images and 526 manually annotated text rectangles. The image column is a Hugging Face image feature containing the original image bytes. The boxes column contains normalized pixel-coordinate bounding boxes in [x_min, y_min, x_max, y_max] format, the original points, label, and shape metadata. annotation_json preserves the… See the full description on the dataset page: https://huggingface.co/datasets/petersunde/manga-covers-text-detection.imageobject-detectionn<1K0 likes57 downloads8d agoHugging Face15UniqueData /ocr-generated-machine-readable-zone-mrz-text-detection OCR GENERATED Machine-Readable Zone (MRZ) Text Detection The dataset includes a collection of GENERATED photos containing Machine Readable Zones (MRZ) commonly found on identification documents such as passports, visas, and ID cards. Each photo in the dataset is accompanied by text detection and Optical Character Recognition (OCR) results. 💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request on our website to… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/ocr-generated-machine-readable-zone-mrz-text-detection.imageimage-to-textn<1K3 likes46 downloads1y agoHugging Face16kevknowscode /ai-text-detection-pile Dataset Card for AI Text Dectection Pile Dataset Summary This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ. Here is the (tentative) breakdown: Human Text Dataset Num Samples Link Reddit WritingPromps 570k Link OpenAI Webtext 260k Link HC3 (Human Responses) 58k Link ivypanda-essays TODO TODO… See the full description on the dataset page: https://huggingface.co/datasets/kevknowscode/ai-text-detection-pile.text1M<n<10M0 likes46 downloads6mo agoHugging Face17Melaraby /EvArEST-dataset-for-Arabic-scene-text-detection EvArEST Everyday Arabic-English Scene Text dataset, from the paper: Arabic Scene Text Recognition in the Deep Learning Era: Analysis on A Novel Dataset Detection Dataset The text detection dataset has 510 images all containing one or more instances of text. Each word is annotated with a four-point polygon that starts with the top left corner of the polygon and follows clockwise. Each image comes with a text file containing three attributes: the four points of the polygon… See the full description on the dataset page: https://huggingface.co/datasets/Melaraby/EvArEST-dataset-for-Arabic-scene-text-detection.imagen<1K0 likes45 downloads2y agoHugging Face18coai /ai-text-detection-benchmarktext1K<n<10K2 likes42 downloads9mo agoHugging Face19sinatras /isogram-ai-text-detection-splits Isogram AI Text Detection Permissive Splits This dataset contains train/validation/test splits for binary AI-generated text detection. It is built from sources whose dataset-level licenses were checked as permissive or public-domain-compatible. Schema text: essay text. label: 0 for human-written text, 1 for AI-generated text. source_dataset: upstream dataset identifier. source_detail: source label retained from the upstream data. source_license: row-level… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/isogram-ai-text-detection-splits.texttext-classification10K<n<100K0 likes38 downloads4mo agoHugging Face20Vxlentina /burmese-text-spam-detection Dataset Card for burmese-text-spam-detection Dataset Description The burmese-text-spam-detection dataset is a high-quality, human-curated collection of 1,000 Burmese text entries specifically designed for binary text classification tasks. The dataset is balanced equally with 500 "spam" and 500 "not_spam" samples. This dataset was compiled to facilitate the development and evaluation of spam-filtering models for the Burmese language, covering diverse sources such… See the full description on the dataset page: https://huggingface.co/datasets/Vxlentina/burmese-text-spam-detection.texttext-classification1K<n<10K0 likes38 downloads4mo agoHugging Face21akgupta4332 /handwritten_text_detection Handwritten text detection dataset Data domain The blanks were provided by youth organization "Armenian Club" (telegram, instagram ), Russia Moscow. The text on blanks was written during dictation "Teladrutyun" in 2018 The blanks were labeled by Amir and Renal during research project in HSE MIEM Dataset info Contains labeled dictations blanks in YOLO format 91 image in total, 73 (80%) for train and 18 (20%) for test No image alignment or any… See the full description on the dataset page: https://huggingface.co/datasets/akgupta4332/handwritten_text_detection.imageobject-detectionn<1K0 likes35 downloads2mo agoHugging Face22optimization-hashira /ai-text-detection-datasettext100K<n<1M0 likes33 downloads1y agoHugging Face23Precious1 /Decoding-Text-Summarization-Most-Frequent-Words-and-Medical-Text-Detectiontextn<1K2 likes19 downloads3y agoHugging Face24kanwal-mehreen18 /Multilingual_Machine_Generated_Text_Detection1 likes19 downloads2y agoHugging Face25R-obi /ai-text-detection-pile-cleaned AI Text Detection Pile - Cleaned Dataset Dataset Description This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training. Dataset Details Total Samples: 721,626 (cleaned from… See the full description on the dataset page: https://huggingface.co/datasets/R-obi/ai-text-detection-pile-cleaned.tabulartext-classification100K<n<1M0 likes19 downloads2mo agoHugging Face26ninaaaaddd /AI_text_detection_dataset0 likes18 downloads3y agoHugging Face27Mharis205 /ai-text-detection-pile-cleaned AI Text Detection Pile - Cleaned Dataset Dataset Description This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training. Dataset Details Total Samples: 721,626 (cleaned from original… See the full description on the dataset page: https://huggingface.co/datasets/Mharis205/ai-text-detection-pile-cleaned.texttext-classification100K<n<1M0 likes15 downloads9mo agoHugging Face28ldiujes /ai_text_detection_dataset_dl_hw_2_v6tabular10K<n<100K0 likes15 downloads6mo agoHugging Face29juno-labs /text-voice-activity-detection license: mit text10K<n<100K0 likes14 downloads1y agoHugging Face30SoyVitou /Khmer-Text-Detection-0.2k SoyVitou/Khmer-Text-Detection-0.2k Khmer OCR dataset for scene text detection + transcription. This dataset is packaged as a Hugging Face dataset using a single Parquet file: train/metadata.parquet ✅ The image column is stored as embedded bytes inside the Parquet, so load_dataset() works without downloading a separate images folder. Dataset format Each row contains: id (string): sample id image (image): image object (decoded by datasets) annotation (string):… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/Khmer-Text-Detection-0.2k.imagen<1K0 likes14 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.