datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ghana-sentences
Ghana Sentences
A growing sentence-level text corpus for Ghanaian languages, tagged with
ISO 639-3 codes and split into per-language subsets. The goal is
broad-coverage text across all Ghanaian languages; this first release draws on
school curriculum materials and a licensing-exam benchmark. More sources will be
added over time.
Language list and ISO codes follow
GhanaNLP/ghana-taught-local-languages.
Loading
from datasets import load_dataset
everything =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-sentences.high-quality-multilingual-sentences
High Quality Multilingual Sentences
This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset.
It includes 1.58 million rows across 51 different languages, each in its own configuration.
Example row (from the all config):
{
"text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.",
"fasttext": "fa",
"gcld3": "fa"
}
Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.sentences
sentences
A multi-register sentence corpus for the Hmar language (hmr, ISO 639-3), containing 260,492 train sentences and 5,316 evaluation sentences (~4.15 million words).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241)
Family: Zo Languages
Volume: 260,492 train sentences | 5,316 evaluation sentences (265,808 total, ~4.15M words)
Validation: 100% verified via hmaraniam… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/sentences.low-quality-multilingual-sentences
Low Quality Multilingual Sentences
This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages.
The new sentences in this dataset are low quality, proceed with caution.
Chinese_sentencesamharic-sentences-corpus
Amharic Sentences Corpus V1.0
Source: Telegram
This 1.6 million Amharic sentences corpus reflects current Amharic usage as of December 20, 2025, and is designed for anyone interested in:
Training Amharic-based LLMs
Fine-tuning NLP models
Building search, summarization, or generative systems in Amharic
The dataset is heavily cleaned and normalized, but like any serious LLM dataset, it still needs proper tokenization for pre-training.
I recommend using an… See the full description on the dataset page: https://huggingface.co/datasets/a3xrfgb/amharic-sentences-corpus.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.English-Hebrew-Mixed-Sentences
English-Hebrew Mixed Sentences Dataset
A dataset of English sentences with Hebrew words and phrases interspersed, designed for speech-to-text training and evaluation for English speakers in Israel.
Overview
This dataset addresses a common challenge for English-speaking immigrants in Israel: standard speech-to-text (STT) systems struggle to accurately transcribe code-switched speech where Hebrew words are mixed into primarily English sentences.
Example: "I need to pick up… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/English-Hebrew-Mixed-Sentences.Yahoo-Finance-News-Sentences
Yahoo Finance News Sentences Dataset
This dataset was created from Yahoo Finance News Articles collected from 6.12.2023 to 20.12.2023
kyrgyz_sentences_with_incorrect_and_correct_umlaut_characters
Kyrgyz Orthographic Correction Dataset
Dataset Description
This dataset is designed to fine-tune language models for a Kyrgyz-to-Kyrgyz orthographic correction task. It addresses a common issue in digital Kyrgyz text where specific Cyrillic characters specific to the Kyrgyz language (ө, ң, ү) are replaced by their Russian keyboard counterparts (о, н, у).
The dataset is structured in a conversational format, making it ideal for instruction-tuning chat models.… See the full description on the dataset page: https://huggingface.co/datasets/murat/kyrgyz_sentences_with_incorrect_and_correct_umlaut_characters.spanglish-sentences
Spanglish Sentences
A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models.
Data format
Each line of spanglish_sentences.jsonl is a JSON object with two fields:
field
description
sentence
A Spanglish utterance (mixed Spanish / English, or monolingual in either language).
english_translation
The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.bonaventure-sentences
Bonaventure on the Sentences (Latin ↔ English)
2,113 chunks of Bonaventure's Commentary on the Sentences and related works, Latin (Quaracchi) aligned with English (~3.33M words). Apparatus (notes, scholia) is kept in a separate field, never merged into the text.
Canonical home: https://bonaventure.wrootpress.com (each record carries its canonical URL). This dataset is a machine-generated export of that site's build, regenerated from the source of truth and never hand-edited; the… See the full description on the dataset page: https://huggingface.co/datasets/wrootpress/bonaventure-sentences.model-card-sentences-annotatedalia_multilingual_parallel_sentences
MULTILINGUAL PARALLEL SENTENCES Dataset
The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models.
It provides aligned sentences in multiple languages to facilitate multilingual learning.
Dataset Structure
The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language.
The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.legalese-sentences_estonian-english
📚 Estonian → English Legal Sentence Translation Dataset
@ajsbsd | Non-commercial Use Only
This dataset contains Estonian legal sentences translated into English, derived from the original paulpall/legalese-sentences_estonian dataset.
Translations were generated using the Helsinki-NLP/opus-mt-et-en model.
🛥️ Dedication
This dataset is dedicated to the pursuit of truth, transparency, and informed discourse in a world shaped by complex global threats.… See the full description on the dataset page: https://huggingface.co/datasets/ajsbsd/legalese-sentences_estonian-english.english_hinglist_sentencesAdaption-multilingual-sentences
This dataset is a remastered version of
Reubencf/PolyglotText
prepared using Adaption's Adaptive Data platform.
Multilingual Sentences (Adaption)
9,999 sentences across 123 languages. A broad multilingual subset of
PolyglotText — originally derived from the
Tatoeba project — with Adaption-sharpened
enhanced_prompt / enhanced_completion / reasoning_trace columns.
Each row carries a source-language sentence, translations, and the
Adaption-processed fields.
Dataset size… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-sentences.bcms-claim-sentencesfa-topic-sentences
README for fa-topic-sentences Dataset
Overview
The fa-topic-sentences dataset is a comprehensive collection of sentences categorized into various topics. Each topic contains approximately 50 sentences in Persian, accompanied by a paraphrased version of each sentence. The dataset is structured in JSON format, providing a straightforward method for accessing individual entries.
Topics Included
The dataset encompasses the following topics:
History
Fashion… See the full description on the dataset page: https://huggingface.co/datasets/mostafaamiri/fa-topic-sentences.cebuano-filipino-sentenceseng-xquad-sentencesrus-xquad-sentencesprocessed-turkish-sentencestech-sentences-error-robustness
Technical English Sentences for NLP Robustness (50k)
Dataset Description
This dataset contains 50,000 synthetically generated English sentences tailored around Software Engineering,
DevOps, and IT contexts.
Key Feature: Intentional Grammatical Anomalies
A unique characteristic of this dataset is the presence of
intentional grammatical and morphological anomalies
in verb forms (e.g., combining past tense with third-person singular endings
like… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/tech-sentences-error-robustness.tbc-sentencesmultilingual-single-sentences
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
multilingual_single_sentences
This dataset consists of single-sentence completions spanning a diverse range of languages, including Russian, German, Spanish, Korean, Chinese, English, French, and Japanese. The content varies widely, covering topics from historical restoration and travel logistics to proverbs and daily observations. Each entry is presented as an isolated text… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-single-sentences.expanded-english-sentences
Expanded English Sentences Dataset
This dataset includes over 15 000 random sentences from the agentlans/high-quality-english-sentences dataset, each paired with a paragraph generated by a customized Llama 3.1 8B model, providing additional context.
Overview
train.jsonl.gz: Contains original sentences and their corresponding AI-generated paragraphs in JSONL (JSON Lines) format compressed using GZip.
Variable
Definition
Type
sentence
Original sentence from the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/expanded-english-sentences.marathi-czech-sentences
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
marathi_czech_sentences
This dataset contains short sentences and questions primarily in Marathi and Czech, covering various conversational contexts. The samples include inquiries about objects, actions, and origins, as well as exclamations and statements. It appears to be a multilingual collection focused on everyday dialogue structures.
Dataset size
There are 3… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/marathi-czech-sentences.uyghur-sentences
🌟 Uyghur AI Corpus: Bridging Heritage & Technology
🌟 ئۇيغۇرچە سۈنئىي ئىدراك خەزىنىسى: مىراس ۋە تېخنىكا كۆۋرۈكى
🌹 Introduction / كىرىش سۆز
In the era of Artificial Intelligence, language is data, and data is survival.
The Uyghur AI Corpus is an initiative to ensure the Uyghur language thrives in the digital age. This dataset serves as a foundational resource to train Large Language Models (LLMs), enabling them to understand, generate… See the full description on the dataset page: https://huggingface.co/datasets/Uyghur-Corpus/uyghur-sentences.icelandic-sentences-gec
