datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Romansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.ParallelThinkingDLMcountdown_problemssosp_sft_datarepro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
bible-parallel-english
Parallel Bible — English Translations and Ancient Versions
A verse-aligned parallel corpus of the Protestant Bible in seventeen English
translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta
for the New Testament.
Looking for every language? This repository is a curated English set,
chosen for spread across translation families and small enough to load whole.
For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses —
see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.screw_retimed_parallel_23_no_fastforward_20260807sango-french-bible-parallel
SFPC: Sango-French Parallel Corpus
The first quality-filtered, verse-aligned Sango-French parallel corpus, constructed for neural machine translation research. This dataset directly addresses the "Sango Problem" identified by Meta's NLLB-200 project — the failure of cross-lingual transfer for a linguistically isolated Creole language.
Associated resources:
Model: alaminerca/nllb-sango-french
Demo: Sango-French Translator
Paper: SangoNMT: Parameter-Efficient Domain Adaptation of… See the full description on the dataset page: https://huggingface.co/datasets/alaminerca/sango-french-bible-parallel.bashkir-russian-parallel
Dataset Card for Bashkir-Russian Parallel Corpus
Dataset Details
Dataset Description
Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart.
The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.vinaya-pitaka-pali-myanmar-parallel
Vinaya Pitaka: Pali-Myanmar Parallel Dataset
Description
This dataset provides a professionally aligned, paragraph-level parallel corpus of the Vinaya Pitaka (The Code of Monastic Discipline). It features the original Pali text (presented in Myanmar script) alongside its modern Myanmar translation.
The dataset covers all five major volumes of the Vinaya:
Pārājika (ပါရာဇိကပါဠိ / ပါရာဇိကဏ်)
Pācittiya (ပါစိတ္တိယပါဠိ / ပါစိတ်)
Mahāvagga (မဟာဝဂ္ဂပါဠိ / မဟာဝါ)
Cūḷavagga… See the full description on the dataset page: https://huggingface.co/datasets/freococo/vinaya-pitaka-pali-myanmar-parallel.ota-bible-parallel
Ottoman–Turkish–English Parallel New Testament Corpus
This repository is a verse-aligned parallel corpus for Ottoman Turkish (Perso-Arabic original script) with English and modern Turkish reference translations.
It is intended for training and evaluating translation models for
Ottoman Turkish, a low-resource historical language.
Source language: Ottoman Turkish (ota), Perso-Arabic script
Target languages: English (en), modern Turkish (tr)
Unit of alignment: a single New… See the full description on the dataset page: https://huggingface.co/datasets/enesyila/ota-bible-parallel.parason-data
parason-data
Evaluation traces and structural-analysis artifacts for the parallel-reasoning line of work.
Training data is not here — it lives in parallel-reasoner/sft-ours (splits 1x, 8x).
This repo holds generated traces, so that structural claims about model behaviour can be
re-derived rather than taken on trust.
Layout
aime24/<model-name>/traces.jsonl
aime24/Qwen3-8B-sft-ours8x-ar/
Traces from parallel-reasoner/Qwen3-8B-sft-ours8x-ar
— the… See the full description on the dataset page: https://huggingface.co/datasets/parallel-reasoner/parason-data.english-classics-parallel-samples
Booklern English classics: parallel samples
Paragraph-aligned opening passages of public-domain English classics with a
translation into Spanish, Japanese, Brazilian Portuguese, Russian, Chinese, published by Booklern, a
bilingual book reader for learning English through real books. Each book is
read on Booklern with a sentence-by-sentence translation under the English,
read-aloud audio, a dictionary and vocabulary tools; the rows here are the
same opening paragraphs that appear… See the full description on the dataset page: https://huggingface.co/datasets/ssergaroo/english-classics-parallel-samples.evenki-rus-parallel-corporaparallel-image-text-dataset-builder
parallel-image-text-dataset-builder (sample)
A small representative sample from the
parallel-image-text-dataset-builder
pipeline: it ingests image-text pairs, removes near-duplicates with
perceptual-hash (dhash) LSH-style bucketing, filters weak pairs by CLIP
image-text similarity, and writes fixed-size WebDataset-style tar shards.
Contents
shard-00002.tar - one WebDataset-style shard (536 samples). Each sample is
two members sharing a key: {key}.jpg (image) and… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/parallel-image-text-dataset-builder.parallel-bm-en
khursanirevo/parallel-bm-en
Parallel English-Bahasa Melayu translation pairs (102k rows, OpenHermes-derived).
Splits
split
rows
train
97,280
validation
5,120
Stratified 95/5 by source/category (seed=42).
Source files
data/sft/parallel_bm_en_30m.jsonl
Schema
Each row is a JSON object. See the loader script for field details.
Provenance
Generated as part of MaLLaM 2026 Tiny pretraining/SFT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/parallel-bm-en.odia-german-parallel-corpus-research
Dataset Summary
This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics.
The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.
