CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AfriSpeech /africa-corpus Africa Corpus Verse-aligned text for 693 African languages, plus several world languages, for building monolingual and parallel corpora. Every language is aligned on a shared verse key, so any single language can be pulled on its own or any two joined into a parallel corpus: Monolingual corpus for any single language African ↔ English (English is the default pair) African ↔ African (e.g. Twi ↔ Yoruba, Hausa ↔ Amharic) African ↔ other language (French, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/africa-corpus.texttranslation10M<n<100M3 likes491 downloads3mo agoHugging Face02SZLHOLDINGS /thesis-corpus-v18 Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. SZLHOLDINGS/thesis-corpus-v18 The v18 Ouroboros Invariant thesis — LaTeX chapters, the 179 formal blocks (theorem / lemma / definition / axiom environments) as a flat CSV, and the per-version delta ledger that tracks how every formal block evolved v1 → v18. Contents File… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/thesis-corpus-v18.texttext-generationn<1K0 likes477 downloads23d agoHugging Face03ghananlpcommunity /ghana-corpus This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana Corpus Verse-aligned text for Ghanaian languages, plus several world languages, for building monolingual and parallel corpora. Every language is aligned on a shared verse key, so any single language can be pulled on its own or… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-corpus.texttranslation1M<n<10M0 likes467 downloads3mo agoHugging Face04LianHong /zomi-monolingual-corpus Zomi Monolingual Corpus v1.0 The Zomi Monolingual Corpus v1.0 contains 363,401 cleaned, deduplicated, reviewed, and permission-approved Zomi sentences. Zomi is represented with the ISO 639-3 language code ctd (Tedim Chin). Quick start from datasets import load_dataset dataset = load_dataset("LianHong/zomi-monolingual-corpus", split="train") print(dataset.num_rows) # 363401 print(dataset[0]["zomi_text"]) Data fields Field Type Description… See the full description on the dataset page: https://huggingface.co/datasets/LianHong/zomi-monolingual-corpus.texttext-generation100K<n<1M0 likes425 downloads26d agoHugging Face05tasksource /blog_authorship_corpustabular100K<n<1M2 likes420 downloads2y agoHugging Face06anonymous-nsc-author /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes319 downloads3mo agoHugging Face07paodigitalhub /blk-text-corpus Verified Pa'O (Blk) Text Corpus - Pa'O Digital Hub Dataset Summary This is the official, verified parallel dataset for the Pa'O language (ISO 639-3: blk) and Burmese (Myanmar) translations, published by Pa'O Digital Hub. The corpus is systematically collected, reviewed, standardized, and verified through the established linguistic and editorial workflow of Pa'O Digital Hub. The Pa'O sentences are based on authentic language usage by Pa'O native speakers and are… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/blk-text-corpus.texttranslationn<1K1 likes306 downloads5d agoHugging Face08yordanoswuletaw /amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus") texttext-generation100M<n<1B4 likes293 downloads2y agoHugging Face09Karavet /ARPA-Armenian-Paraphrase-Corpus Dataset Description We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language. Dataset Summary The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.text1K<n<10K3 likes290 downloads4y agoHugging Face10SolarisCipher /hk_content_corpus HK Content Corpus (Cantonese & Traditional Chinese) This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms. It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling. Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators. This… See the full description on the dataset page: https://huggingface.co/datasets/SolarisCipher/hk_content_corpus.text1M<n<10M0 likes269 downloads1y agoHugging Face11kritsadaK /EDGAR-CORPUS-Financial-Summarization EDGAR-CORPUS : 10K Financial Report Summarization Extracted from SEC EDGAR filings (1993-2020). This dataset enhances financial report summarization by leveraging a hybrid AI model strategy. Using: ChatGPT-3.5 Turbo(~70%), Claude 3.5 (~30% to generate structured, accurate, and concise summaries) Dataset Composition Summaries in this dataset are generated using a hybrid AI model strategy, balancing quality and efficiency:ChatGPT-3.5 Turbo (~70%) – Used for structured… See the full description on the dataset page: https://huggingface.co/datasets/kritsadaK/EDGAR-CORPUS-Financial-Summarization.textsummarization10K<n<100K4 likes253 downloads2y agoHugging Face12zeroshot /cybersecurity-corpustext1K<n<10K10 likes252 downloads3y agoHugging Face13myothiha /mm_eng_alt_corpustext10K<n<100K0 likes215 downloads3y agoHugging Face14bekan /english_karakalpak_parallel_corpus_v5 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa). Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.texttranslation10K<n<100K4 likes210 downloads9d agoHugging Face15mariagrandury /fake_news_corpus_spanish Fake News Corpus Spanish Citation Gómez-Adorno, H., Posadas-Durán, J. P., Enguix, G. B., & Capetillo, C. P. (2021). Overview of FakeDeS at IberLEF 2021: Fake News Detection in Spanish Shared Task. Procesamiento del Lenguaje Natural, 67, 223-231. Aragón, M. E., Jarquín, H., Gómez, M. M. Y., Escalante, H. J., Villaseñor-Pineda, L., Gómez-Adorno, H., ... & Posadas-Durán, J. P. (2020, September). Overview of mex-a3t at iberlef 2020: Fake news and aggressiveness analysis in… See the full description on the dataset page: https://huggingface.co/datasets/mariagrandury/fake_news_corpus_spanish.texttext-classificationn<1K2 likes206 downloads2y agoHugging Face16talkmap /banking-conversation-corpus Banking 300k Dataset Overview This dataset consists of 300,000 synthetically generated conversations in a customer service setting for the telecom industry. There are two speakers: a customer, and an agent. texttext-generation1M<n<10M17 likes203 downloads3y agoHugging Face17CtnkyaABC /turkish-law-corpus ⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA 🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır. 🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.textquestion-answering10K<n<100K3 likes188 downloads2mo agoHugging Face18nawabhussain /Kashmiri-Language-Corpus Kashmiri Textual Data Corpus Introduction This repository contains a combined dataset of Kashmiri textual data collected from various sources. The data has been sourced from different locations and may contain non-Kashmiri text (e.g., Urdu, Persian). The goal of this corpus is to provide a wide variety of Kashmiri text data for research and language processing tasks. Sources of Data 1. mzmmoazam/kashmiri_dataset (HTML Data) Source: GitHub -… See the full description on the dataset page: https://huggingface.co/datasets/nawabhussain/Kashmiri-Language-Corpus.text10K<n<100K2 likes185 downloads2y agoHugging Face19allegro /summarization-polish-summaries-corpustext10K<n<100K5 likes180 downloads5y agoHugging Face20tasal9 /Pashto-Textbooks-PDFs-Corpus Pashto Textbooks and PDFs Corpus Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tabulartext-generationn<1K0 likes179 downloads2mo agoHugging Face21matthh /gutenberg-poetry-corpustabular1K<n<10K7 likes153 downloads4y agoHugging Face22lhbelfanti /drug-use-corpus Drug Use Corpus (Spanish) Binary classification dataset for drug use detection in Spanish tweets Drug Use Corpus (Spanish) This dataset contains Spanish-language tweets related to drug use, specifically focusing on references to marijuana, cocaine, and other substances. The dataset is designed for binary classification tasks in the context of substance use detection in social media discussions. Dataset Description The Drug Use Corpus consists of 3,000… See the full description on the dataset page: https://huggingface.co/datasets/lhbelfanti/drug-use-corpus.texttext-classification10K<n<100K0 likes150 downloads28d agoHugging Face23anthonyyazdaniml /gliner-biomed-curated-corpus GLiNER-BioMed curated corpus Unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @misc{yazdani2025glinerbiomedsuiteefficientmodels, title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition}, author={Anthony Yazdani and Ihor Stepanov and Douglas Teodoro}… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-curated-corpus.text100K<n<1M0 likes142 downloads1y agoHugging Face24sarnab /Shakespeare_Corpustext1K<n<10K1 likes138 downloads3y agoHugging Face25drelhaj /Arabic-news-and-management-corpus Arabic Management, Economics & Financial News Corpus (1,200 Articles) This corpus contains 1,200 Arabic news and management articles drawn from three distinct domains. It was originally compiled as part of research into Arabic Corpus Linguistics, management communication, financial discourse and domain-specific NLP. Both plain text and POS-tagged versions are available. The dataset has been widely used in teaching and research, including the King Saud University book Corpus… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-news-and-management-corpus.texttext-classification1K<n<10K0 likes130 downloads10mo agoHugging Face26mihai-chindris /policy-rag-corpus-metadata Policy RAG Corpus Metadata (No Raw Data) This repository is a metadata-only companion for the Policy RAG project built for the Quantic MSSE AI Engineering program. It does not include the actual PDF files. The source PDFs are hosted in the companion GitHub repository. What this repo includes metadata.csv: structured metadata for 11 policy documents (filename, title, category, page count, source type, description) Citation and provenance notes for reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/mihai-chindris/policy-rag-corpus-metadata.textquestion-answeringn<1K1 likes129 downloads3mo agoHugging Face27talkmap /telecom-conversation-corpus Telecom 200k Dataset Overview This dataset consists of 200,000 synthetically generated conversations in a customer service setting for the telecom industry. There are two speakers: a customer, and an agent. texttext-generation1M<n<10M22 likes126 downloads3y agoHugging Face28fliarbi /urban-heat-research-corpus Urban Heat Research Corpus (UHRC) v1.0 What does the world study, invent and report about urban heat? This dataset puts three records of the same problem side by side: 20,422 research papers on urban heat islands and extreme heat in cities (1990–2025) with the claims their abstracts make, 106,458 news articles about heat (2021–2025) coded for 51 subjects, framings and terms, and 4,123 patent families for heat-mitigation technologies (2006–2024) — plus supplementary tables on the… See the full description on the dataset page: https://huggingface.co/datasets/fliarbi/urban-heat-research-corpus.imagetext-classification100K<n<1M0 likes126 downloads4d agoHugging Face29Okwu /african-language-parallel-corpus African Language Parallel Corpus Human-created, human-validated parallel sentence pairs for three African languages, released openly by Okwu. Version 1.0. Dataset summary A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own language-learning curriculum — content authored and reviewed by native-speaker educators — supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.texttranslation10K<n<100K0 likes121 downloads8h agoHugging Face30anthonyyazdaniml /gliner-biomed-balanced-curated-corpus GLiNER-BioMed balanced curated corpus Balanced, unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @misc{yazdani2025glinerbiomedsuiteefficientmodels, title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition}, author={Anthony Yazdani and Ihor Stepanov and Douglas… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-balanced-curated-corpus.text100K<n<1M0 likes114 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.