CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shangrilar /ko_text2sqltext10K<n<100K19 likes1.6k downloads3y agoHugging Face02GesturingMan /CPC_Text_Roughtext1K<n<10K0 likes937 downloads3y agoHugging Face03CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes816 downloads5mo agoHugging Face04Duyacquy /Ecommerce_texttext10K<n<100K0 likes795 downloads1y agoHugging Face05omar-sharif /BAD-Bengali-Aggressive-Text-Dataset Novel Aggressive Text Dataset in Bengali Tackling Cyber-Aggression: Identification and Fine-Grained Categorization of Aggressive Texts on Social Media using Weighted Ensemble of Transformers Author: Omar Sharif and Mohammed Moshiul Hoque Related Papers: Paper1 in Neurocomputing Journal Paper2 in CONSTRAINT@AAAI-2021 Paper3 in LTEDI@EACL-2021 Abstract The pervasiveness of aggressive content in social media has become a serious concern for government… See the full description on the dataset page: https://huggingface.co/datasets/omar-sharif/BAD-Bengali-Aggressive-Text-Dataset.tabular10K<n<100K3 likes645 downloads5y agoHugging Face06FredZhang7 /toxi-text-3MThis is a large multilingual toxicity dataset with 3M rows of text data from 55 natural languages, all of which are written/sent by humans, not machine translation models. The preprocessed training data alone consists of 2,880,667 rows of comments, tweets, and messages. Among these rows, 416,529 are classified as toxic, while the remaining 2,463,773 are considered neutral. Below is a table to illustrate the data composition: Toxic Neutral Total multilingual-train-deduplicated.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/toxi-text-3M.texttext-classification1M<n<10M32 likes631 downloads1y agoHugging Face07dmitva /human_ai_generated_text Human or AI-Generated Text The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems. File Name model_training_dataset.csv File Structure id: Unique identifier for each record. human_text: Human-written content. ai_text: AI-generated texts. instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.text1M<n<10M38 likes526 downloads3y agoHugging Face08neurovlm /embedded_texttextn<1K0 likes379 downloads4mo agoHugging Face09shibing624 /chinese_text_correction Dataset Card 中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。 Repository: shibing624/pycorrector Dataset Summary 拼写纠错数据 lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2 ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.text100K<n<1M15 likes332 downloads2y agoHugging Face10ThomasTheMaker /TextToCadQuery-1text100K<n<1M0 likes321 downloads1y agoHugging Face11paodigitalhub /blk-text-corpus Verified Pa'O (Blk) Text Corpus - Pa'O Digital Hub Dataset Summary This is the official, verified parallel dataset for the Pa'O language (ISO 639-3: blk) and Burmese (Myanmar) translations, published by Pa'O Digital Hub. The corpus is systematically collected, reviewed, standardized, and verified through the established linguistic and editorial workflow of Pa'O Digital Hub. The Pa'O sentences are based on authentic language usage by Pa'O native speakers and are… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/blk-text-corpus.texttranslationn<1K1 likes306 downloads5d agoHugging Face12ChaseLabs /Harmful-Texts-On-Mastodon 🦣 Mastodon Wild Data for Harmful Content Detection Overview The Harmful Texts on Mastodon dataset is a human-annotated corpus of 3,000 English posts collected from the decentralized social media platform Mastodon between December 2024 and February 2025.It is designed to evaluate the robustness, generalization, and personalization capabilities of large language models (LLMs) and in-context learning (ICL) approaches for harmful content detection in real-world scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/ChaseLabs/Harmful-Texts-On-Mastodon.texttext-classification1K<n<10K2 likes301 downloads11mo agoHugging Face13Ateeqq /AI-and-Human-Generated-Text AI & Human Generated Text I am Using this dataset for AI Text Detection for https://exnrt.com. Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA Description The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.texttext-classification10K<n<100K24 likes280 downloads2y agoHugging Face14LaconicAI /text_message_function_calling_open_chatThis is a small synthetic dataset to model a function call for text messaging someone from a cell phone. This has been tested with and used to finetune a set of smaller models and deployed directly on the pixel 8 pro and Fold 4 phones. texttext-generation10K<n<100K2 likes258 downloads2y agoHugging Face15Kazimir-ai /text-to-image-prompts The dataset of the most popular text-to-image prompts. Dataset Details Dataset Description Curated by: kazimir.ai Funded by [optional]: [More Information Needed] Shared by [optional]: https://kazimir.ai License: apache-2.0 Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Free to use. Dataset Structure CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Kazimir-ai/text-to-image-prompts.text10K<n<100K9 likes254 downloads3y agoHugging Face16Navanjana /ARCHIVE-TEXT-URLS Internet Archive English Text URLs Dataset Dataset Description This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials. Dataset Summary Total Rows: 11,151,637 Language: English Source: Internet Archive Format: CSV with metadata and direct text file URLs Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.texttext-generation1M<n<10M1 likes244 downloads10mo agoHugging Face17simana /textclassificationMNLItext100K<n<1M0 likes235 downloads4y agoHugging Face18owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K28 likes221 downloads4y agoHugging Face19zxbsmk /laion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs). image10M<n<100M1 likes219 downloads3y agoHugging Face20smit18 /text_emotiontext10K<n<100K1 likes218 downloads3y agoHugging Face21ImruQays /Rasaif-Classical-Arabic-English-Parallel-texts Introduction This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period. Content Details Contained within this dataset are English translations of the following texts, sourced from the Rasaif website: A Muslim Manual of War Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.texttranslation10K<n<100K8 likes207 downloads3y agoHugging Face22silentone0725 /ai-human-text-detection-v1 🧠 AI vs Human Text Detection Dataset (v1) This dataset merges nine major public and academic corpora to form one of the most comprehensive resources for AI-generated text detection model training and evaluation. 🔗 Sources The dataset consolidates, cleans, and standardizes multiple open datasets and research benchmarks, each focusing on human vs. AI-generated text classification: Hello-SimpleAI / HC3 — Human–ChatGPT comparison corpus gsingh1-py / train — Large-scale… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/ai-human-text-detection-v1.text10K<n<100K8 likes198 downloads11mo agoHugging Face23tasal9 /Pashto-Textbooks-PDFs-Corpus Pashto Textbooks and PDFs Corpus Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tabulartext-generationn<1K0 likes179 downloads2mo agoHugging Face24knowledgator /Scientific-text-classificationtext10K<n<100K17 likes167 downloads3y agoHugging Face25Ghana-NLP /ENGLISH_TWI_PARALLEL_TEXT GhanaNLP Twi and English Parallel Data Twi_to_English • 1 MB • XLS English_to_Twi • 1 MB • XLS The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.text1K<n<10K3 likes148 downloads10mo agoHugging Face26Krooz /Campus_Recruitment_Text Dataset Description This data set consists of Placement data of students in a XYZ campus. Based on the student's performance report we are classifying his Placement Status. The dataset is derived from a csv data. The Mistral7B model is used with data-to-text methodology to convert each of the rows in the csv data into a textual format for the LLM's, the conversion script is in this notebook. The Prompt field is the prompt used on Mistral7B LLM and the response field is the… See the full description on the dataset page: https://huggingface.co/datasets/Krooz/Campus_Recruitment_Text.text10K<n<100K0 likes145 downloads2y agoHugging Face27md-nishat-008 /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.texttext-generation10K<n<100K2 likes135 downloads1y agoHugging Face28Faith1712 /allsides_text_proper_truncatedtext1K<n<10K0 likes134 downloads2y agoHugging Face29Yah216 /Poem_APCD_text_onlyWe used the APCD dataset cited hereafter for pretraining the model. The dataset has been cleaned and only the main text column was kept: @Article{Yousef2019LearningMetersArabicEnglish-arxiv, author = {Yousef, Waleed A. and Ibrahime, Omar M. and Madbouly, Taha M. and Mahmoud, Moustafa A.}, title = {Learning Meters of Arabic and English Poems With Recurrent Neural Networks: a Step Forward for Language Understanding and Synthesis}, journal =… See the full description on the dataset page: https://huggingface.co/datasets/Yah216/Poem_APCD_text_only.text1M<n<10M0 likes132 downloads4y agoHugging Face30used255 /youtube_annotations_text Youtube Annotations Text YouTube 注释(YouTube Annotations)是 YouTube 在 2008 年推出的一项功能, 允许视频创作者在视频上添加文本、链接和互动元素, 以增强观众的观看体验. YouTube 已在 2019 年删除了此功能. 您可以在这里找到由 omarroth 创建的存档 YouTube Annotations, 本数据集从13亿条存档中提取出了文本. 如果您需要 x_id 与 videoId 的映射, 请使用 utilities/video_text_mapping_indexed.sqlite3 数据库. text10M<n<100M1 likes126 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.