CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01s-nlp /paradetox ParaDetox: Text Detoxification with Parallel Data (English) This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.imagetext-generation10K<n<100K9 likes901 downloads1y agoHugging Face02ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K2 likes223 downloads3y agoHugging Face03s-nlp /ru_paradetox ParaDetox: Text Detoxification with Parallel Data (Russian) This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit [2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.imagetext-generation10K<n<100K4 likes209 downloads1y agoHugging Face04jpwahle /machine-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools. It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses). The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.texttext-classification100K<n<1M7 likes111 downloads1y agoHugging Face05jpwahle /autoencoder-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models. It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses). The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.tabulartext-classification1M<n<10M2 likes87 downloads1y agoHugging Face06abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes57 downloads3mo agoHugging Face07ClarusC64 /clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1 Goal Test if a model can hold separate reasoning streams at once Detect constraint dismissal Detect bleed-over where one stream turns into claims in the other What it measures streams_heldResponse acknowledges and maintains both streams bleed_overConstraint stream improperly becomes a medical claim, or vice versa premature_synthesisResponse forces a single solution that silences one stream assumption_collapseResponse drops a premise entirely Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.texttext-generationn<1K0 likes44 downloads8mo agoHugging Face08jpwahle /autoregressive-paraphrase-dataset Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset.texttext-classification100K<n<1M1 likes42 downloads4y agoHugging Face09DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes40 downloads4mo agoHugging Face10faisal4590aziz /bangla-health-related-paraphrased-dataset Dataset Card for "BanglaHealthParaphrase" BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.tabulartext-generation100K<n<1M2 likes32 downloads1y agoHugging Face11ReliableAI /Irish-English-Parallel-Collection UCCIX's English-Irish Parallel Textual Corpus Dataset Summary This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR. This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data. Dataset Sources Source Description Statistics Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.texttext-generation10K<n<100K1 likes30 downloads2y agoHugging Face12abdelhaqueidali /Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset Dataset Card for Tachelhit Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Tachelhit language ($\text{Tacelḥit}$ / $\text{Tamazigt}$), pairing native Latin-based orthography with automated Amazigh script transliterations. It is built by processing clean source sentences through an algorithmic engine designed to handle phonetic mappings, manage contextual schwa distributions, and safeguard acronyms and foreign proper names. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset.texttranslation10K<n<100K0 likes30 downloads4mo agoHugging Face13ParasiticRogue /Fullmoon-LightFinalized version of the Bluemoon-Light dataset. Fully trimmed, cleaned, and grammar checked three times over: First by me in ridding it of obvious unwanted junk, second by an AI to grammer/spell chack it and other fixes such as adding in quotes where the dialogue had none, and then finally by me again to make sure the AI didn't add it's own junk back in. The dataset has been edited for better parquet quantization such as exl2 or gguf, making models slightly more stable during creative… See the full description on the dataset page: https://huggingface.co/datasets/ParasiticRogue/Fullmoon-Light.texttext-generation1K<n<10K5 likes25 downloads2y agoHugging Face14SuryaKrishna02 /aya-telugu-paraphrase Summary aya-telugu-paraphrase is an open source dataset of instruct-style records generated from the Telugu split of ai4bharat/IndicXParaphrase dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-paraphrase.texttext-generation1K<n<10K4 likes18 downloads3y agoHugging Face15s-nlp /paranmt_for_detox ParaNMTDetox: Detoxification with Parallel Data (English) This repository contains information about filtered ParaNMT dataset for text detoxification task. Here, we have paraphrasing pairs where one text is toxic and another is non-toxic. Toxicity levels were defined by English toxicity classifier. The original paper "ParaDetox: Detoxification with Parallel Data" with SOTA text detoxification was presented at ACL 2022 main conference. ParaNMTDetox Filtering Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paranmt_for_detox.texttext-generation1K<n<10K0 likes17 downloads3y agoHugging Face16prithivMLmods /GPT-Paraphrases GPT-Paraphrases dataset This dataset contains text passages and their paraphrases generated using the GPT-3 language model. The paraphrases are designed to be semantically equivalent to the original text, but with different wording and structure. The dataset includes text formatted in JSON and is in English. Dataset Statistics Number of text passages: Not specified in the information you provided. Source of text passages: Not specified in the information you provided.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/GPT-Paraphrases.texttext-generation100K<n<1M1 likes15 downloads2y agoHugging Face17DiligentPenguinn /vietnamese-author-styles-paraphrasedtexttext-generationn<1K0 likes12 downloads1y agoHugging Face18shaheeeeeeeeem /parabank-paraphrase-labeled ParaBank Paraphrase Labeled (Style-Conditioned) 200,001-row style-labeled paraphrase dataset derived from ParaBank v2, balanced at 66,667 rows per style: [FORMAL], [CASUAL] (informal), [POLITE]. Columns reference: source English sentence paraphrase_1: best-ranked paraphrase from ParaBank v2 style_label: 0=polite, 1=formal, 2=informal plus raw formality/politeness classifier labels and confidence scores Data Attribution Derived from ParaBank 2: J.… See the full description on the dataset page: https://huggingface.co/datasets/shaheeeeeeeeem/parabank-paraphrase-labeled.tabulartext-generation100K<n<1M0 likes12 downloads3mo agoHugging Face19ramsleeb /paradetox ParaDetox: Text Detoxification with Parallel Data (English) This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/ramsleeb/paradetox.imagetext-generation10K<n<100K0 likes11 downloads7mo agoHugging Face20bekan /english_karakalpak_parallel_corpus_v1 English-Karakalpak Parallel Corpus (en-kaa) Dataset Description English-Karakalpak Parallel Corpus is a high-quality dataset containing 10,441 aligned sentence pairs in English and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. The corpus utilizes the official… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v1.texttranslation10K<n<100K2 likes9 downloads10mo agoHugging Face21ParasiticRogue /Bluemoon-vicunaOld version. Use the cleaned and grammar checked version below: https://huggingface.co/datasets/ParasiticRogue/Bluemoon-Light texttext-generationn<1K2 likes5 downloads3y agoHugging Face22bekan /english_karakalpak_pairs_parallel_corpus_v2_8907 English-Karakalpak Parallel Corpus v2 (8.9K) Dataset Description English-Karakalpak Parallel Corpus v2 is a high-quality dataset containing 8,906 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. This resource… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_pairs_parallel_corpus_v2_8907.texttranslation1K<n<10K1 likes5 downloads10mo agoHugging Face23Dddixyy /Italian_latin_parallel_animals descrizioni di animali e habitat - Synthetic Dataset This dataset was generated using the Synthetic Dataset Generator powered by Gemini AI. Topic: descrizioni di animali e habitat Field 1: italiano Field 2: latino antico(traduzione) Rows: 280 Generated on: 2025-05-27T00:07:49.042Z texttext-generationn<1K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.