CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artem9k /ai-text-detection-pile Dataset Card for AI Text Dectection Pile Dataset Summary This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ. Here is the (tentative) breakdown: Human Text Dataset Num Samples Link Reddit WritingPromps 570k Link OpenAI Webtext 260k Link HC3 (Human Responses) 58k Link ivypanda-essays TODO TODO… See the full description on the dataset page: https://huggingface.co/datasets/artem9k/ai-text-detection-pile.text1M<n<10M46 likes557 downloads4y agoHugging Face02coai /ai-text-detection-trainingtext10K<n<100K3 likes158 downloads9mo agoHugging Face03srikanthgali /ai-text-detection-pile-cleaned AI Text Detection Pile - Cleaned Dataset Dataset Description This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training. Dataset Details Total Samples: 721,626 (cleaned from original… See the full description on the dataset page: https://huggingface.co/datasets/srikanthgali/ai-text-detection-pile-cleaned.texttext-classification100K<n<1M5 likes124 downloads1y agoHugging Face04amaye15 /receipts-text-detectionimage10K<n<100K3 likes102 downloads2y agoHugging Face05Mharis205 /ai-text-detection-pile Dataset Card for AI Text Dectection Pile Dataset Summary This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ. Here is the (tentative) breakdown: Human Text Dataset Num Samples Link Reddit WritingPromps 570k Link OpenAI Webtext 260k Link HC3 (Human Responses) 58k Link ivypanda-essays TODO TODO… See the full description on the dataset page: https://huggingface.co/datasets/Mharis205/ai-text-detection-pile.text1M<n<10M0 likes88 downloads9mo agoHugging Face06petersunde /manga-covers-text-detection Manga Covers Text Detection Manually annotated text boxes for manga covers text detection created by @JustANormalTinkerer. This dataset contains 100 cover images and 526 manually annotated text rectangles. The image column is a Hugging Face image feature containing the original image bytes. The boxes column contains normalized pixel-coordinate bounding boxes in [x_min, y_min, x_max, y_max] format, the original points, label, and shape metadata. annotation_json preserves the… See the full description on the dataset page: https://huggingface.co/datasets/petersunde/manga-covers-text-detection.imageobject-detectionn<1K0 likes60 downloads12d agoHugging Face07kevknowscode /ai-text-detection-pile Dataset Card for AI Text Dectection Pile Dataset Summary This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ. Here is the (tentative) breakdown: Human Text Dataset Num Samples Link Reddit WritingPromps 570k Link OpenAI Webtext 260k Link HC3 (Human Responses) 58k Link ivypanda-essays TODO TODO… See the full description on the dataset page: https://huggingface.co/datasets/kevknowscode/ai-text-detection-pile.text1M<n<10M0 likes55 downloads6mo agoHugging Face08Melaraby /EvArEST-dataset-for-Arabic-scene-text-detection EvArEST Everyday Arabic-English Scene Text dataset, from the paper: Arabic Scene Text Recognition in the Deep Learning Era: Analysis on A Novel Dataset Detection Dataset The text detection dataset has 510 images all containing one or more instances of text. Each word is annotated with a four-point polygon that starts with the top left corner of the polygon and follows clockwise. Each image comes with a text file containing three attributes: the four points of the polygon… See the full description on the dataset page: https://huggingface.co/datasets/Melaraby/EvArEST-dataset-for-Arabic-scene-text-detection.imagen<1K0 likes53 downloads2y agoHugging Face09sinatras /isogram-ai-text-detection-splits Isogram AI Text Detection Permissive Splits This dataset contains train/validation/test splits for binary AI-generated text detection. It is built from sources whose dataset-level licenses were checked as permissive or public-domain-compatible. Schema text: essay text. label: 0 for human-written text, 1 for AI-generated text. source_dataset: upstream dataset identifier. source_detail: source label retained from the upstream data. source_license: row-level… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/isogram-ai-text-detection-splits.texttext-classification10K<n<100K0 likes35 downloads4mo agoHugging Face10optimization-hashira /ai-text-detection-datasettext100K<n<1M0 likes33 downloads1y agoHugging Face11coai /ai-text-detection-benchmarktext1K<n<10K2 likes32 downloads9mo agoHugging Face12R-obi /ai-text-detection-pile-cleaned AI Text Detection Pile - Cleaned Dataset Dataset Description This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training. Dataset Details Total Samples: 721,626 (cleaned from… See the full description on the dataset page: https://huggingface.co/datasets/R-obi/ai-text-detection-pile-cleaned.tabulartext-classification100K<n<1M0 likes24 downloads3mo agoHugging Face13ldiujes /ai_text_detection_dataset_dl_hw_2_v6tabular10K<n<100K0 likes19 downloads6mo agoHugging Face14Mharis205 /ai-text-detection-pile-cleaned AI Text Detection Pile - Cleaned Dataset Dataset Description This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training. Dataset Details Total Samples: 721,626 (cleaned from original… See the full description on the dataset page: https://huggingface.co/datasets/Mharis205/ai-text-detection-pile-cleaned.texttext-classification100K<n<1M0 likes15 downloads9mo agoHugging Face15SoyVitou /Khmer-Text-Detection-0.2k SoyVitou/Khmer-Text-Detection-0.2k Khmer OCR dataset for scene text detection + transcription. This dataset is packaged as a Hugging Face dataset using a single Parquet file: train/metadata.parquet ✅ The image column is stored as embedded bytes inside the Parquet, so load_dataset() works without downloading a separate images folder. Dataset format Each row contains: id (string): sample id image (image): image object (decoded by datasets) annotation (string):… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/Khmer-Text-Detection-0.2k.imagen<1K0 likes14 downloads7mo agoHugging Face16kahua-ml /experimento3-industrial-text-detection Experimento-3 - Industrial Machinery Text Detection Dataset Dataset Description This dataset contains 4,237 images of industrial machinery nameplates with detailed text field annotations for OCR and information extraction tasks. The dataset focuses on extracting key information from equipment nameplates including manufacturer, model, serial numbers, and dates. Dataset Summary Task: Industrial text detection and OCR Domain: Industrial machinery and equipment… See the full description on the dataset page: https://huggingface.co/datasets/kahua-ml/experimento3-industrial-text-detection.image1K<n<10K0 likes13 downloads1y agoHugging Face17author-ai-text-detect /longlamp-ai-detectiongated author-ai-text-detect/longlamp-ai-detection A fork of LongLaMP/LongLaMP (user-setting configs) with model completions attached. Every row is a LongLaMP sample with its original fields intact — the author field, input, output and the full profile — plus a completions list holding the model-written texts for that same task input. The human reference is output; each entry of completions is machine-written. There is no label column because the nesting already says which is which.… See the full description on the dataset page: https://huggingface.co/datasets/author-ai-text-detect/longlamp-ai-detection.texttext-classification1K<n<10K0 likes13 downloads3d agoHugging Face18juno-labs /text-voice-activity-detection license: mit text10K<n<100K0 likes12 downloads1y agoHugging Face19Varun53 /AI_text_detectiontext1K<n<10K2 likes10 downloads3y agoHugging Face20gunnybd01 /ai-text-detection-progressivetext100K<n<1M0 likes9 downloads5mo agoHugging Face21joshuakelleych /text_detections_easyocr_yolov10imagen<1K1 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.