CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Columbia-NLP /PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles. Code: https://github.com/siyan-sylvia-li/PAPILLON texttext-generationn<1K3 likes6.7k downloads1y agoHugging Face02s-nlp /paradetox ParaDetox: Text Detoxification with Parallel Data (English) This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.imagetext-generation10K<n<100K9 likes843 downloads1y agoHugging Face03s-nlp /ru_paradetox ParaDetox: Text Detoxification with Parallel Data (Russian) This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit [2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.imagetext-generation10K<n<100K4 likes212 downloads1y agoHugging Face04Finnish-NLP /mc4_fi_cleaned Dataset Card for mC4 Finnish Cleaned Dataset Summary mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split. Supported Tasks and Leaderboards mC4 Finnish is mainly intended to pretrain Finnish language models and word representations. Languages Finnish Dataset Structure Data Instances [Needs More Information] Data Fields The data have several fields: url: url of the source as a string text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.texttext-generation10M<n<100M4 likes203 downloads4y agoHugging Face05upb-nlp /RoJBMO RoJBMO: Junior Balkan Mathematical Olympiad Benchmark RoJBMO is a benchmark of 508 competition mathematics problems drawn from the Junior Balkan Mathematical Olympiad (JBMO), its official shortlists, and the Romanian Team Selection Tests (TST). It is designed to evaluate the mathematical reasoning capabilities of large language models on problems that are underrepresented in existing benchmarks and resistant to training data contamination. Sources Source… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/RoJBMO.tabulartext-generationn<1K1 likes87 downloads19d agoHugging Face06biu-nlp /sentence_union_generation Revisiting Sentence Union Generation as a Testbed for Text Consolidation Eran Hirsch1, Valentina Pyatkin1, Ruben Wolhandler1, Avi Caciularu1, Asi Shefer2, Ido Dagan1 1Bar-Ilan University, 2One AI This is the official dataset of the paper "Revisiting Sentence Union Generation as a Testbed for Text Consolidation". Paper 📄 (Findings of ACL 2023) Code 💻 Abstract Tasks involving text generation based on multiple input texts, such as multi-document summarization… See the full description on the dataset page: https://huggingface.co/datasets/biu-nlp/sentence_union_generation.texttext-generation1K<n<10K3 likes64 downloads3y agoHugging Face07DAMO-NLP-SG /SOULThis repo contains the data for our paper "SOUL: Towards Sentiment and Opinion Understanding of Language" in EMNLP 2023. Github repo Statistics The SOUL dataset comprises 15,028 statements related to 3,638 reviews, resulting in an average of 4.13 statements per review. To create training, development, and test sets, we split the reviews in a ratio of 6:1:3, respectively. Split # reviews # statements True False Not-given Total Train 2,182 3,675 2,159 8,834 3,000 8,834… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/SOUL.texttext-classification10K<n<100K0 likes57 downloads3y agoHugging Face08NLPinas /ph_en_text_detoxedPhEnText Detoxed is a large-scale and multi-domain lexical data written in Philippine English and Taglish text. The news articles, religious articles and court decisions collated by the original researchers were filtered for toxicity and special characters were further preprocessed. This dataset has been configured to easily fine-tune LLaMA-based models (Alpaca, Guanaco, Vicuna, LLaMA 2, etc.) In total, this dataset contains 6.29 million rows of training data and 2.7 million rows of testing… See the full description on the dataset page: https://huggingface.co/datasets/NLPinas/ph_en_text_detoxed.texttext-generation1M<n<10M2 likes55 downloads3y agoHugging Face09NLPForUA /dumy-zno-ukrainian-math-history-geo-r1-o1 DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers) DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks. The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian: Думи мої, думи мої, Лихо мені з вами! Нащо стали на папері Сумними рядами?.. Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.tabulartext-generation1K<n<10K2 likes32 downloads1y agoHugging Face10talhaa /nlp-google-reviews-dataset NLP Google Reviews Dataset A curated, multi-source dataset of 516 real Google reviews prepared for NLP tasks such as sentiment analysis, text classification, and topic modelling. Dataset Description This dataset was built using a production-grade Python pipeline that collects Google reviews from three independent sources, cleans and normalizes the data, and merges everything into a single structured CSV. Sources public_dataset: 495 reviews web_scraping: 16… See the full description on the dataset page: https://huggingface.co/datasets/talhaa/nlp-google-reviews-dataset.texttext-classificationn<1K0 likes30 downloads4mo agoHugging Face11ud-nlp /LLM-Text-Generation-Dataset Generated Text Dataset - 4 Millions+ Logs Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data Dataset characteristics: Characteristic Data Description Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.texttext-generation1K<n<10K0 likes23 downloads1y agoHugging Face12propfirmkey /trading-finance-glossary-nlp Trading & Finance Glossary Dataset for NLP A comprehensive, structured dataset of 209 trading and financial terminology entries designed for natural language processing applications in the financial domain. Description This dataset provides a curated collection of trading and financial terms with rich metadata including definitions, categorical labels, semantic relationships, contextual usage examples, and difficulty classifications. Each entry has been written to reflect… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/trading-finance-glossary-nlp.texttext-classificationn<1K2 likes19 downloads6mo agoHugging Face13IndoHealth-NLP /PubMed-Bilingual-Medical-Sample-EN-ID 🏥 PubMed Bilingual Medical Sample (EN-ID) Providing premium, high-quality English-Indonesian bilingual medical datasets for AI, NLP, and Machine Learning research. 📥 Download Free Sample You can directly access and download the free dataset sample (.csv format) from our repository files here: ⬇️ Download Free Sample File (tree/main) 🚀 Upgrade to Full Version (Volume 1) This repository contains a free sample of our industry-grade parallel… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/PubMed-Bilingual-Medical-Sample-EN-ID.texttranslationn<1K0 likes17 downloads2mo agoHugging Face14s-nlp /paranmt_for_detox ParaNMTDetox: Detoxification with Parallel Data (English) This repository contains information about filtered ParaNMT dataset for text detoxification task. Here, we have paraphrasing pairs where one text is toxic and another is non-toxic. Toxicity levels were defined by English toxicity classifier. The original paper "ParaDetox: Detoxification with Parallel Data" with SOTA text detoxification was presented at ACL 2022 main conference. ParaNMTDetox Filtering Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paranmt_for_detox.texttext-generation1K<n<10K0 likes15 downloads3y agoHugging Face15NeuraXenetica /managpt-4080-nlp-prompts-and-generated-textsThis dataset includes 4,080 texts that were generated by the ManaGPT-1020 large language model, in response to particular input sequences. ManaGPT-1020 is a free, open-source model available for download and use via Hugging Face’s “transformers” Python package. The model is a 1.5-billion-parameter LLM that’s capable of generating text in order to complete a sentence whose first words have been provided via a user-supplied input sequence. The model represents an elaboration of GPT-2 that has… See the full description on the dataset page: https://huggingface.co/datasets/NeuraXenetica/managpt-4080-nlp-prompts-and-generated-texts.texttext-generation1K<n<10K1 likes14 downloads3y agoHugging Face16NLPForUA /ua-code-benchgated LLM Code Generation Benchmark for Ukrainian language Preprint: https://arxiv.org/pdf/2511.05040 Updates 17/10/2025: paper presented at "Informatics. Culture. Technology" conference; 18/09/2025: added data preparation and evaluation notebooks (check notebooks readme first); 17/09/2025: updated result chart; added gpt-5, gpt-oss, and grok-4 evaluations. Thousands of programming tasks in Ukrainian language combined with graded Python solutions (code… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/ua-code-bench.tabulartext-generation1K<n<10K0 likes11 downloads2mo agoHugging Face17NLPC-UOM /anonymized-sinhala-letter-corpus Anonymized Sinhala Official Letter Corpus A small, hand-curated corpus of 151 formal Sinhala letters, fully anonymized with bracketed placeholders. It is intended for training and evaluating models that generate, complete, or classify Sinhala official correspondence — a task with very little public training data. Dataset at a glance Examples 151 Language Sinhala (si) Register Formal throughout Letter length 42–240 words (median 108, mean 114)… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/anonymized-sinhala-letter-corpus.texttext-generationn<1K0 likes9 downloads2mo agoHugging Face18sinhala-nlp /NSINA-Headlinesgated Sinhala Headline Generation This is a text generation task created with the NSINA dataset. This dataset is also released with the same license as NSINA. The objective of the task is to generate news headlines based on the provided news content. Data We used the same instances from NSINA 1.0 as all the news articles had headlines. We divided this dataset into a training and test set following a 0.8 split. Data can be loaded into pandas dataframes using the following code.… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA-Headlines.texttext-generation100K<n<1M0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.