datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-paragraphs
wikipedia-paragraphs
wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research.
Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID.
Dataset structure
Configurations
The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.parallel-sentences-ccmatrix
Dataset Card for Parallel Sentences - CCMatrix
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The texts originate from the CCMatrix dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse
parallel-sentences-jw300
parallel-sentences-news-commentary… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-ccmatrix.st-parallel-sentences
Dataset Card for "st-parallel-sentences"
More Information needed
NCERT-Parallel-Dataset-Indicparacrawl_context
Dataset Card for ParaCrawl_Context
This is a dataset for document-level machine translation introduced in the ACL 2024 paper Document-Level Machine Translation with Large-Scale Public Parallel Data. It is a dataset consisting of parallel sentence pairs from the ParaCrawl dataset along with corresponding preceding context extracted from the webpages the sentences were crawled from.
Dataset Details
Dataset Description
This dataset adds document-level… See the full description on the dataset page: https://huggingface.co/datasets/Proyag/paracrawl_context.parallel-sentences-talks
Dataset Card for Parallel Sentences - Talks
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the Talks dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-talks.kasem-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Kasem Speech-Text Parallel Dataset
Dataset Description
This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.parallel-sentences-opensubtitles
Dataset Card for Parallel Sentences - OpenSubtitles
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the OpenSubtitles dataset.
Warning! The quality of this dataset is not great; many of the english and non-english texts don't match well, or are fully empty.
Related Datasets
The following… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-opensubtitles.parallel-sentences-jw300
Dataset Card for Parallel Sentences - JW300
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the JW300 dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-jw300.xone-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..)
A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/xone-repository-parallel-en-id-corpus.parallel-sentences-wikimatrix
Dataset Card for Parallel Sentences - WikiMatrix
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the WikiMatrix dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-wikimatrix.parallel-sentences-opus-100
Dataset Card for Parallel Sentences - OPUS-100
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The sentences originate from the OPUS-100 website.
In particular, this dataset is a reformatting of the OPUS-100 dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-opus-100.opus_paracrawl
Dataset Card for OpusParaCrawl
Dataset Summary
Parallel corpora from Web Crawls collected in the ParaCrawl project.
Tha dataset contains:
42 languages, 43 bitexts
total number of files: 59,996
total number of tokens: 56.11G
total number of sentence fragments: 3.13G
To load a language pair which isn't part of the config, all you need to do is specify the language code as pairs,
e.g.
dataset = load_dataset("opus_paracrawl", lang1="en", lang2="so")
You can find the valid… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/opus_paracrawl.polynews-parallel
Dataset Card for PolyNewsParallel
Dataset Summary
PolyNewsParallel is a multilingual paralllel dataset containing news titles for 833 language pairs. It covers 64 languages and 17 scripts.
Uses
This dataset can be used for machine translation or text retrieval.
Languages
There are 64 languages avaiable:
Code
Language
Script
amh_Ethi
Amharic
Ethiopic
arb_Arab
Modern Standard Arabic
Arabic
ayr_Latn
Central Aymara
Latin
bam_Latn
Bambara… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews-parallel.ga-speech-text-parallel-90k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Speech-Text Parallel Dataset
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.twi-trigrams-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.parallel-sentences-europarl
Dataset Card for Parallel Sentences - Europarl
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the Europarl dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices
parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-europarl.gemstones_data_order_parallelGemstones Training Dataset - Parallel workers sharded version
This data is a reprocessed version of the first 1B rows of the Dolma v1.7 dataset (https://huggingface.co/datasets/allenai/dolma).
The data is encoded using the Pythia tokenizer: https://huggingface.co/EleutherAI/pythia-160m
Disclaimer: this is an approximation of the dataset used to train the Gemstones model suite.
Due to the randomized and sharded nature of the distributed training code, the only way to perfectly
reproduce the… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/gemstones_data_order_parallel.ccel-paragraphs
CCEL Paragraphs
Dataset Description
Dataset Summary
This dataset includes all paragraphs from the Christian Classics Ethereal Library. It also includes scripture references extracted from the ThML.
Supported Tasks and Leaderboards
It is expected that this dataset can be used as part of the training pipeline for large language models. In particular, it could be used to create a clustering benchmark by using scripture references as labels.… See the full description on the dataset page: https://huggingface.co/datasets/jncraton/ccel-paragraphs.robotwin_put_obj_cabinet_parallel_50
ABORTED research artifact — retained for reproducibility, not deleted.
This artifact is retained as historical evidence only. Its cached segment-start main-camera observations are not eligible for dynamic-main-view claims. See ABORTED.yaml for the machine-readable archival record.
Archival registry mapping:
experiment_id: E002
This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha"… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_put_obj_cabinet_parallel_50.parallel-sentences-global-voices
Dataset Card for Parallel Sentences - Global Voices
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the Global Voices dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-global-voices.jitteredwebsites-merged-224-paraphrasedX-ALMA-Parallel-Data
This is the translation parallel dataset used by X-ALMA.
@misc{xu2024xalmaplugplay,
title={X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale},
author={Haoran Xu and Kenton Murray and Philipp Koehn and Hieu Hoang and Akiko Eriguchi and Huda Khayrallah},
year={2024},
eprint={2410.03115},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.03115},
}
Franka_robotiq_85_parallel_simThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda_roboriq",
"total_episodes": 4060,
"total_frames": 709512,
"total_tasks": 31,
"total_videos": 0,
"total_chunks": 5,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:4060"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyzhang01/Franka_robotiq_85_parallel_sim.multilingual-wikipedia-paragraphshuman-ai-parallel-corpus
Human-AI Parallel English Corpus (HAP-E) 🙃
Purpose
The HAP-E corpus is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs).
The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words.
Thus, a second 500-word chunk of human-authored text (what actually comes next in the original text) can be compared to… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus.ALMA-Human-Parallel
Dataset Card for "ALMA-Human-Parallel"
This is human-written parallel dataset used by ALMA translation models.
@misc{xu2023paradigm,
title={A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models},
author={Haoran Xu and Young Jin Kim and Amr Sharaf and Hany Hassan Awadalla},
year={2023},
eprint={2309.11674},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
@misc{xu2024contrastive,
title={Contrastive… See the full description on the dataset page: https://huggingface.co/datasets/haoranxu/ALMA-Human-Parallel.IISC_parallel_franka_all_fix_tracks_fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 15,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyzhang01/IISC_parallel_franka_all_fix_tracks_fixed.common_voice_16_1_hi_pseudo_labelled4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the
exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
