datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ghana-sentences
Ghana Sentences
A growing sentence-level text corpus for Ghanaian languages, tagged with
ISO 639-3 codes and split into per-language subsets. The goal is
broad-coverage text across all Ghanaian languages; this first release draws on
school curriculum materials and a licensing-exam benchmark. More sources will be
added over time.
Language list and ISO codes follow
GhanaNLP/ghana-taught-local-languages.
Loading
from datasets import load_dataset
everything =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-sentences.high-quality-multilingual-sentences
High Quality Multilingual Sentences
This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset.
It includes 1.58 million rows across 51 different languages, each in its own configuration.
Example row (from the all config):
{
"text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.",
"fasttext": "fa",
"gcld3": "fa"
}
Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.low-quality-multilingual-sentences
Low Quality Multilingual Sentences
This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages.
The new sentences in this dataset are low quality, proceed with caution.
spanglish-sentences
Spanglish Sentences
A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models.
Data format
Each line of spanglish_sentences.jsonl is a JSON object with two fields:
field
description
sentence
A Spanglish utterance (mixed Spanish / English, or monolingual in either language).
english_translation
The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.alia_multilingual_parallel_sentences
MULTILINGUAL PARALLEL SENTENCES Dataset
The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models.
It provides aligned sentences in multiple languages to facilitate multilingual learning.
Dataset Structure
The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language.
The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.Adaption-multilingual-sentences
This dataset is a remastered version of
Reubencf/PolyglotText
prepared using Adaption's Adaptive Data platform.
Multilingual Sentences (Adaption)
9,999 sentences across 123 languages. A broad multilingual subset of
PolyglotText — originally derived from the
Tatoeba project — with Adaption-sharpened
enhanced_prompt / enhanced_completion / reasoning_trace columns.
Each row carries a source-language sentence, translations, and the
Adaption-processed fields.
Dataset size… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-sentences.expanded-english-sentences
Expanded English Sentences Dataset
This dataset includes over 15 000 random sentences from the agentlans/high-quality-english-sentences dataset, each paired with a paragraph generated by a customized Llama 3.1 8B model, providing additional context.
Overview
train.jsonl.gz: Contains original sentences and their corresponding AI-generated paragraphs in JSONL (JSON Lines) format compressed using GZip.
Variable
Definition
Type
sentence
Original sentence from the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/expanded-english-sentences.uyghur-sentences
🌟 Uyghur AI Corpus: Bridging Heritage & Technology
🌟 ئۇيغۇرچە سۈنئىي ئىدراك خەزىنىسى: مىراس ۋە تېخنىكا كۆۋرۈكى
🌹 Introduction / كىرىش سۆز
In the era of Artificial Intelligence, language is data, and data is survival.
The Uyghur AI Corpus is an initiative to ensure the Uyghur language thrives in the digital age. This dataset serves as a foundational resource to train Large Language Models (LLMs), enabling them to understand, generate… See the full description on the dataset page: https://huggingface.co/datasets/Uyghur-Corpus/uyghur-sentences.Taiwan-Text-Excellence-sentences
台灣文摘句資料集
概述
Taiwan-Text-Excellence 句子資料集是從較大的 liswei/Taiwan-Text-Excellence-2B 資料集中抽取的 200 萬個獨特中文句子的綜合集。這些句子是隨機選取的,並使用 chinese-sentence-processor 工具進行分割。此資料集非常適合各種自然語言處理任務,包括語言建模、文本生成和其他研究用途。
資料集統計
總句數: 2,000,000
訓練集: 1,600,000 個句子
測試集: 400,000 個句子
資料格式
資料集中的每一行都包含一個欄位:
**text**:包含中文句子的字串。
範例
{"text": "而新郎和女方家人的脂燭在當晚亦會合二為一,再送到母屋帳前點燃一個燈籠,保持三日不滅。"}
{"text": "這個時期簽署的現代劇至今仍是台灣戲劇的中堅力量,而這十年為之奮鬥也奠定了其後數十年的基礎。"}
{"text": "曾柏瑜今天也車票,明天還請在規劃畫相關票活動中。"}… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/Taiwan-Text-Excellence-sentences.
