datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parallel-sentences-ccmatrix
Dataset Card for Parallel Sentences - CCMatrix
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The texts originate from the CCMatrix dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse
parallel-sentences-jw300
parallel-sentences-news-commentary… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-ccmatrix.st-parallel-sentences
Dataset Card for "st-parallel-sentences"
More Information needed
parallel-sentences-talks
Dataset Card for Parallel Sentences - Talks
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the Talks dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-talks.parallel-sentences-opensubtitles
Dataset Card for Parallel Sentences - OpenSubtitles
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the OpenSubtitles dataset.
Warning! The quality of this dataset is not great; many of the english and non-english texts don't match well, or are fully empty.
Related Datasets
The following… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-opensubtitles.parallel-sentences-jw300
Dataset Card for Parallel Sentences - JW300
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the JW300 dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-jw300.parallel-sentences-tatoeba
Dataset Card for Parallel Sentences - Tatoeba
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the Tatoeba dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-tatoeba.parallel-sentences-opus-100
Dataset Card for Parallel Sentences - OPUS-100
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The sentences originate from the OPUS-100 website.
In particular, this dataset is a reformatting of the OPUS-100 dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-opus-100.parallel-sentences-wikimatrix
Dataset Card for Parallel Sentences - WikiMatrix
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the WikiMatrix dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-wikimatrix.tamil_sentences_master_raw
Dataset Card for "tamil_sentences_master"
More Information needed
parallel-sentences-europarl
Dataset Card for Parallel Sentences - Europarl
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the Europarl dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-europarl.bias-test-gpt-sentences
Dataset Card for "BiasTestGPT: Generated Test Sentences"
Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models.
This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool.
BiasTestGPT HuggingFace Tool
Dataset with Bias Specifications
Project Landing Page
Dataset Structure
The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.CCCPT-splited_preprocessed_max1024sz_sentenceshigh-quality-english-sentences
High-Quality English Sentences
Dataset Description
This dataset contains a collection of high-quality English sentences sourced from C4 and FineWeb (not FineWeb-Edu). The sentences have been carefully filtered and processed to ensure quality and uniqueness.
"High-quality" means they're legible English and not spam, although they may still have spelling and grammar errors.
Source Data
Before filtering:
C4: 1 million sentences
FineWeb: 1 million sentences… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-english-sentences.parallel-sentences-global-voices
Dataset Card for Parallel Sentences - Global Voices
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the Global Voices dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-global-voices.ghana-sentences
Ghana Sentences
A growing sentence-level text corpus for Ghanaian languages, tagged with
ISO 639-3 codes and split into per-language subsets. The goal is
broad-coverage text across all Ghanaian languages; this first release draws on
school curriculum materials and a licensing-exam benchmark. More sources will be
added over time.
Language list and ISO codes follow
GhanaNLP/ghana-taught-local-languages.
Loading
from datasets import load_dataset
everything =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-sentences.hsk-sentences-audio
HSK Sentences Audio
4,354 Chinese sentences graded against the official HSK 3.0 levels 1–6, with
pinyin, English translations, per-word glosses, grammar tags, and normal/slow
synthetic speech. The complete export contains 8,708 MP3 files.
Dataset structure
The Viewer reads native Parquet from data/train.parquet, avoiding a dependency
on Hugging Face's JSON-to-Parquet conversion service. The same 4,354 records are
also available as validated JSON Lines in… See the full description on the dataset page: https://huggingface.co/datasets/no7z/hsk-sentences-audio.SAE_activations_modal_sentencesparallel-sentences-news-commentary
Dataset Card for Parallel Sentences - News Commentary
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the News-Commentary dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-news-commentary.exp01-eeg-to-text-sentences
Exp01 — Sentence-level EEG-to-text training data (unified)
This is a private working corpus for experiment 1 (fine-tuning EEG / time-series
foundation models on EEG-to-English-text). It bundles several public EEG-while-reading
datasets into a single, raw-lossless parquet schema where one row = one sentence read by
one participant.
⚠️ License: Per-source licenses are preserved verbatim in each row's license
column and source_url. Do not re-distribute publicly without re-checking the… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/exp01-eeg-to-text-sentences.financial_phrasebank_sentences_allagree
Dataset Card for financial_phrasebank
Dataset Summary
Polar sentiment dataset of sentences from financial news. The dataset consists of 4840 sentences from English language financial news categorised by sentiment. The dataset is divided by agreement rate of 5-8 annotators.
Supported Tasks and Leaderboards
Sentiment Classification
Languages
English
Dataset Structure
Data Instances
{ "sentence": "Pharmaceuticals group Orion Corp… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/financial_phrasebank_sentences_allagree.high-quality-multilingual-sentences
High Quality Multilingual Sentences
This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset.
It includes 1.58 million rows across 51 different languages, each in its own configuration.
Example row (from the all config):
{
"text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.",
"fasttext": "fa",
"gcld3": "fa"
}
Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.fineweb-young-sentencesbias-test-gpt-sentencesTaiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.zinc-sentencesml_parallel_sentences_250kmultilingual-sentences
Multilingual Sentences
Dataset contains sentences from 50 languages, grouped by their two-letter ISO 639-1 codes. The "all" configuration includes sentences from all languages.
Dataset Overview
Multilingual Sentence Dataset is a comprehensive collection of high-quality, linguistically diverse sentences. Dataset is designed to support a wide range of natural language processing tasks, including but not limited to language modeling, machine translation, and cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-sentences.yoda_sentences
Yoda Speak
This small dataset was built using two resources:
Harvard Sentences, a list of 720 short sentences grouped into 72 sets of 10 sentences each
English to Yoda Translator, an online translator that converts normal English into Yoda's way of speaking.
Fun with this dataset I hope you have! Yes, hrrrm.
wikipedia-en-sentences
Dataset Card for Wikipedia Sentences (English)
This dataset contains 7.87 million English sentences and can be used in knowledge distillation of embedding models.
Dataset Details
Columns: "sentence"
Column types: str
Examples:{
'sentence': "After the deal was approved and NONG's stock rose to $13, Farris purchased 10,000 shares at the $2.50 price, sold 2,500 shares at the new price to reimburse the company, and gave the remaining 7,500 shares to Landreville at no cost… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/wikipedia-en-sentences.dhivehi-noisy-sentences
Dhivehi Noisy Sentences Dataset
This dataset contains parallel examples of clean text and text with introduced errors across three categories: spelling, grammar, and punctuation.
Dataset Description
This dataset is designed to train models that can correct errors in Dhivehi text. Each example consists of:
clean_text: The correct, error-free Dhivehi text
noisy_text: The same text with introduced errors
error_type: The category of error (spelling, grammar, or punctuation)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-noisy-sentences.
