datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles.
Code: https://github.com/siyan-sylvia-li/PAPILLON
paradetox
ParaDetox: Text Detoxification with Parallel Data (English)
This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.ru_paradetox
ParaDetox: Text Detoxification with Parallel Data (Russian)
This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit
[2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.mc4_fi_cleaned
Dataset Card for mC4 Finnish Cleaned
Dataset Summary
mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split.
Supported Tasks and Leaderboards
mC4 Finnish is mainly intended to pretrain Finnish language models and word representations.
Languages
Finnish
Dataset Structure
Data Instances
[Needs More Information]
Data Fields
The data have several fields:
url: url of the source as a string
text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.RoJBMO
RoJBMO: Junior Balkan Mathematical Olympiad Benchmark
RoJBMO is a benchmark of 508 competition mathematics problems drawn from the Junior Balkan Mathematical Olympiad (JBMO), its official shortlists, and the Romanian Team Selection Tests (TST). It is designed to evaluate the mathematical reasoning capabilities of large language models on problems that are underrepresented in existing benchmarks and resistant to training data contamination.
Sources
Source… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/RoJBMO.sentence_union_generation
Revisiting Sentence Union Generation as a Testbed for Text Consolidation
Eran Hirsch1,
Valentina Pyatkin1,
Ruben Wolhandler1,
Avi Caciularu1,
Asi Shefer2,
Ido Dagan1
1Bar-Ilan University, 2One AI
This is the official dataset of the paper "Revisiting Sentence Union Generation as a Testbed for Text Consolidation".
Paper 📄 (Findings of ACL 2023)
Code 💻
Abstract
Tasks involving text generation based on multiple input texts, such as multi-document summarization… See the full description on the dataset page: https://huggingface.co/datasets/biu-nlp/sentence_union_generation.SOULThis repo contains the data for our paper "SOUL: Towards Sentiment and Opinion Understanding of Language" in EMNLP 2023.
Github repo
Statistics
The SOUL dataset comprises 15,028 statements related to 3,638 reviews, resulting in an average of 4.13 statements per review. To create training, development, and test sets, we split the reviews in a ratio of 6:1:3, respectively.
Split
# reviews
# statements
True
False
Not-given
Total
Train
2,182
3,675
2,159
8,834
3,000
8,834… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/SOUL.ph_en_text_detoxedPhEnText Detoxed is a large-scale and multi-domain lexical data written in Philippine English and Taglish text. The news articles, religious articles and court decisions collated by the original researchers were filtered for toxicity and special characters were further preprocessed. This dataset has been configured to easily fine-tune LLaMA-based models (Alpaca, Guanaco, Vicuna, LLaMA 2, etc.) In total, this dataset contains 6.29 million rows of training data and 2.7 million rows of testing… See the full description on the dataset page: https://huggingface.co/datasets/NLPinas/ph_en_text_detoxed.dumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.nlp-google-reviews-dataset
NLP Google Reviews Dataset
A curated, multi-source dataset of 516 real Google reviews prepared for NLP tasks such as sentiment analysis, text classification, and topic modelling.
Dataset Description
This dataset was built using a production-grade Python pipeline that collects Google reviews from three independent sources, cleans and normalizes the data, and merges everything into a single structured CSV.
Sources
public_dataset: 495 reviews
web_scraping: 16… See the full description on the dataset page: https://huggingface.co/datasets/talhaa/nlp-google-reviews-dataset.LLM-Text-Generation-Dataset
Generated Text Dataset - 4 Millions+ Logs
Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data
Dataset characteristics:
Characteristic
Data
Description
Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.trading-finance-glossary-nlp
Trading & Finance Glossary Dataset for NLP
A comprehensive, structured dataset of 209 trading and financial terminology entries designed for natural language processing applications in the financial domain.
Description
This dataset provides a curated collection of trading and financial terms with rich metadata including definitions, categorical labels, semantic relationships, contextual usage examples, and difficulty classifications. Each entry has been written to reflect… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/trading-finance-glossary-nlp.PubMed-Bilingual-Medical-Sample-EN-ID
🏥 PubMed Bilingual Medical Sample (EN-ID)
Providing premium, high-quality English-Indonesian bilingual medical datasets for AI, NLP, and Machine Learning research.
📥 Download Free Sample
You can directly access and download the free dataset sample (.csv format) from our repository files here:
⬇️ Download Free Sample File (tree/main)
🚀 Upgrade to Full Version (Volume 1)
This repository contains a free sample of our industry-grade parallel… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/PubMed-Bilingual-Medical-Sample-EN-ID.paranmt_for_detox
ParaNMTDetox: Detoxification with Parallel Data (English)
This repository contains information about filtered ParaNMT dataset for text detoxification task. Here, we have paraphrasing pairs where one text is toxic and another is non-toxic. Toxicity levels were defined by English toxicity classifier.
The original paper "ParaDetox: Detoxification with Parallel Data" with SOTA text detoxification was presented at ACL 2022 main conference.
ParaNMTDetox Filtering Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paranmt_for_detox.managpt-4080-nlp-prompts-and-generated-textsThis dataset includes 4,080 texts that were generated by the ManaGPT-1020 large language model, in response to particular input sequences.
ManaGPT-1020 is a free, open-source model available for download and use via Hugging Face’s “transformers” Python package. The model is a 1.5-billion-parameter LLM that’s capable of generating text in order to complete a sentence whose first words have been provided via a user-supplied input sequence. The model represents an elaboration of GPT-2 that has… See the full description on the dataset page: https://huggingface.co/datasets/NeuraXenetica/managpt-4080-nlp-prompts-and-generated-texts.ua-code-bench
LLM Code Generation Benchmark for Ukrainian language
Preprint: https://arxiv.org/pdf/2511.05040
Updates
17/10/2025: paper presented at "Informatics. Culture. Technology" conference;
18/09/2025: added data preparation and evaluation notebooks (check notebooks readme first);
17/09/2025: updated result chart; added gpt-5, gpt-oss, and grok-4 evaluations.
Thousands of programming tasks in Ukrainian language combined with graded Python solutions (code… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/ua-code-bench.anonymized-sinhala-letter-corpus
Anonymized Sinhala Official Letter Corpus
A small, hand-curated corpus of 151 formal Sinhala letters, fully anonymized with
bracketed placeholders. It is intended for training and evaluating models that generate,
complete, or classify Sinhala official correspondence — a task with very little public
training data.
Dataset at a glance
Examples
151
Language
Sinhala (si)
Register
Formal throughout
Letter length
42–240 words (median 108, mean 114)… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/anonymized-sinhala-letter-corpus.NSINA-Headlines
Sinhala Headline Generation
This is a text generation task created with the NSINA dataset. This dataset is also released with the same license as NSINA. The objective of the task is to generate news headlines based on the provided news content.
Data
We used the same instances from NSINA 1.0 as all the news articles had headlines. We divided this dataset into a training and test set following a 0.8 split.
Data can be loaded into pandas dataframes using the following code.… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA-Headlines.
