datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
africa-corpus
Africa Corpus
Verse-aligned text for 693 African languages, plus several world languages,
for building monolingual and parallel corpora. Every language is aligned
on a shared verse key, so any single language can be pulled on its own or any
two joined into a parallel corpus:
Monolingual corpus for any single language
African ↔ English (English is the default pair)
African ↔ African (e.g. Twi ↔ Yoruba, Hausa ↔ Amharic)
African ↔ other language (French, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/africa-corpus.thesis-corpus-v18
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/thesis-corpus-v18
The v18 Ouroboros Invariant thesis — LaTeX chapters, the 179 formal blocks
(theorem / lemma / definition / axiom environments) as a flat CSV, and the per-version
delta ledger that tracks how every formal block evolved v1 → v18.
Contents
File… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/thesis-corpus-v18.ghana-corpus
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Corpus
Verse-aligned text for Ghanaian languages, plus several world languages, for
building monolingual and parallel corpora. Every language is aligned on a
shared verse key, so any single language can be pulled on its own or… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-corpus.zomi-monolingual-corpus
Zomi Monolingual Corpus v1.0
The Zomi Monolingual Corpus v1.0 contains 363,401 cleaned, deduplicated,
reviewed, and permission-approved Zomi sentences. Zomi is represented with the
ISO 639-3 language code ctd (Tedim Chin).
Quick start
from datasets import load_dataset
dataset = load_dataset("LianHong/zomi-monolingual-corpus", split="train")
print(dataset.num_rows) # 363401
print(dataset[0]["zomi_text"])
Data fields
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/LianHong/zomi-monolingual-corpus.blog_authorship_corpusNeapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.blk-text-corpus
Verified Pa'O (Blk) Text Corpus - Pa'O Digital Hub
Dataset Summary
This is the official, verified parallel dataset for the Pa'O language (ISO 639-3: blk) and Burmese (Myanmar) translations, published by Pa'O Digital Hub.
The corpus is systematically collected, reviewed, standardized, and verified through the established linguistic and editorial workflow of Pa'O Digital Hub. The Pa'O sentences are based on authentic language usage by Pa'O native speakers and are… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/blk-text-corpus.amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic.
You can load the dataset as follows
from datasets import load_dataset
ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus")
ARPA-Armenian-Paraphrase-Corpus
Dataset Description
We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language.
Dataset Summary
The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.hk_content_corpus
HK Content Corpus (Cantonese & Traditional Chinese)
This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms.
It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling.
Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators.
This… See the full description on the dataset page: https://huggingface.co/datasets/SolarisCipher/hk_content_corpus.EDGAR-CORPUS-Financial-Summarization
EDGAR-CORPUS : 10K Financial Report Summarization
Extracted from SEC EDGAR filings (1993-2020). This dataset enhances financial report summarization by leveraging a hybrid AI model strategy.
Using:
ChatGPT-3.5 Turbo(~70%),
Claude 3.5 (~30% to generate structured, accurate, and concise summaries)
Dataset Composition
Summaries in this dataset are generated using a hybrid AI model strategy, balancing quality and efficiency:ChatGPT-3.5 Turbo (~70%) – Used for structured… See the full description on the dataset page: https://huggingface.co/datasets/kritsadaK/EDGAR-CORPUS-Financial-Summarization.cybersecurity-corpusmm_eng_alt_corpusenglish_karakalpak_parallel_corpus_v5
English-Karakalpak Parallel Corpus
This dataset contains parallel sentences in English and Karakalpak language.
It is created to support AI development for the Karakalpak language.
Dataset Description
English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.fake_news_corpus_spanish
Fake News Corpus Spanish
Citation
Gómez-Adorno, H., Posadas-Durán, J. P., Enguix, G. B., & Capetillo, C. P. (2021). Overview of FakeDeS at IberLEF 2021: Fake News Detection in Spanish Shared Task. Procesamiento del Lenguaje Natural, 67, 223-231.
Aragón, M. E., Jarquín, H., Gómez, M. M. Y., Escalante, H. J., Villaseñor-Pineda, L., Gómez-Adorno, H., ... & Posadas-Durán, J. P. (2020, September). Overview of mex-a3t at iberlef 2020: Fake news and aggressiveness analysis in… See the full description on the dataset page: https://huggingface.co/datasets/mariagrandury/fake_news_corpus_spanish.banking-conversation-corpus
Banking 300k Dataset Overview
This dataset consists of 300,000 synthetically generated conversations in a customer service setting for the telecom industry. There are two speakers: a customer, and an agent.
turkish-law-corpus
⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA
🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır.
🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.Kashmiri-Language-Corpus
Kashmiri Textual Data Corpus
Introduction
This repository contains a combined dataset of Kashmiri textual data collected from various sources. The data has been sourced from different locations and may contain non-Kashmiri text (e.g., Urdu, Persian). The goal of this corpus is to provide a wide variety of Kashmiri text data for research and language processing tasks.
Sources of Data
1. mzmmoazam/kashmiri_dataset (HTML Data)
Source: GitHub -… See the full description on the dataset page: https://huggingface.co/datasets/nawabhussain/Kashmiri-Language-Corpus.summarization-polish-summaries-corpusPashto-Textbooks-PDFs-Corpus
Pashto Textbooks and PDFs Corpus
Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.gutenberg-poetry-corpusdrug-use-corpus
Drug Use Corpus (Spanish)
Binary classification dataset for drug use detection in Spanish tweets
Drug Use Corpus (Spanish)
This dataset contains Spanish-language tweets related to drug use, specifically focusing on references to marijuana, cocaine, and other substances. The dataset is designed for binary classification tasks in the context of substance use detection in social media discussions.
Dataset Description
The Drug Use Corpus consists of 3,000… See the full description on the dataset page: https://huggingface.co/datasets/lhbelfanti/drug-use-corpus.gliner-biomed-curated-corpus
GLiNER-BioMed curated corpus
Unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@misc{yazdani2025glinerbiomedsuiteefficientmodels,
title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition},
author={Anthony Yazdani and Ihor Stepanov and Douglas Teodoro}… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-curated-corpus.Shakespeare_CorpusArabic-news-and-management-corpus
Arabic Management, Economics & Financial News Corpus (1,200 Articles)
This corpus contains 1,200 Arabic news and management articles drawn from three distinct domains. It was originally compiled as part of research into Arabic Corpus Linguistics, management communication, financial discourse and domain-specific NLP. Both plain text and POS-tagged versions are available.
The dataset has been widely used in teaching and research, including the King Saud University book Corpus… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-news-and-management-corpus.policy-rag-corpus-metadata
Policy RAG Corpus Metadata (No Raw Data)
This repository is a metadata-only companion for the Policy RAG project built for the Quantic MSSE AI Engineering program.
It does not include the actual PDF files. The source PDFs are hosted in the companion GitHub repository.
What this repo includes
metadata.csv: structured metadata for 11 policy documents (filename, title, category, page count, source type, description)
Citation and provenance notes for reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/mihai-chindris/policy-rag-corpus-metadata.telecom-conversation-corpus
Telecom 200k Dataset Overview
This dataset consists of 200,000 synthetically generated conversations in a customer service setting for the telecom industry. There are two speakers: a customer, and an agent.
urban-heat-research-corpus
Urban Heat Research Corpus (UHRC) v1.0
What does the world study, invent and report about urban heat? This dataset puts three records of the same problem side by side:
20,422 research papers on urban heat islands and extreme heat in cities (1990–2025) with the claims their abstracts make,
106,458 news articles about heat (2021–2025) coded for 51 subjects, framings and terms, and
4,123 patent families for heat-mitigation technologies (2006–2024) — plus supplementary tables on the… See the full description on the dataset page: https://huggingface.co/datasets/fliarbi/urban-heat-research-corpus.african-language-parallel-corpus
African Language Parallel Corpus
Human-created, human-validated parallel sentence pairs for three African languages,
released openly by Okwu. Version 1.0.
Dataset summary
A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and
Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own
language-learning curriculum — content authored and reviewed by native-speaker educators —
supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.gliner-biomed-balanced-curated-corpus
GLiNER-BioMed balanced curated corpus
Balanced, unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@misc{yazdani2025glinerbiomedsuiteefficientmodels,
title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition},
author={Anthony Yazdani and Ihor Stepanov and Douglas… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-balanced-curated-corpus.
