datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GSM8KInstruct_ParallelCyber-Parallel-Dataset-IndicFinance-Parallel-Dataset-IndicParallel_Dataset
Dataset Sources
Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation
Link: https://aclanthology.org/2025.findings-acl.1200.pdf
Repository: https://github.com/CONE-MT/CONE
Law-Parallel-Dataset-IndicRomansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.Medical-Parallel-Dataset-IndicParallelThinkingDLMCA-Parallel-Dataset-Indictrilingual-parallel-phrasebooks-bgpu
Bashkir Trilingual Parallel Phrasebooks
9,857 phrases aligned across three languages — Bashkir, Russian and one of Altai, Arabic, Kazakh, Yakut (Sakha), Chinese — from five phrasebooks published by M. Akmulla Bashkir State Pedagogical University. One row is one phrase in all three languages: a parallel corpus for machine translation and cross-lingual work with a low-resource Turkic language. Each phrasebook is a separate file and a separate config, because the third language… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/trilingual-parallel-phrasebooks-bgpu.cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences.
BFCL-V4-Parallel-Native
BFCL V4 Parallel Native
Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
Categories
live_parallel
live_parallel_multiple
parallel
parallel_multiple
Counts
train: 352 rows
eval: 88 rows
total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.parallel_ab-ru
Dataset Summary
The Abkhaz Russian parallel corpus dataset is a collection of 205,665 sentences/words extracted from different sources; e-books, web scrapping.
Dataset Creation
Source Data
Here is a link to the source on github
Considerations for Using the Data
Other Known Limitations
The accuracy of the dataset is around 95% (gramatical, arthographical errors)
Jee-Parallel-Dataset-IndicEgyptian-Arabic-English-Parallel-Corpus
Egyptian Arabic-English Parallel Corpus
Author: Mohamed Abdalkader · LinkedIn · GitHub
A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation.
Dataset Structure
egyptian-arabic-english-parallel-corpus/
├── SFT/
│ ├── Train/
│ │ ├── topics/ # 1,800 individual topic JSON files
│ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.ParallelFiction-Ja_En-100k
Dataset details:
Each entry in this dataset is a sentence-aligned Japanese web novel chapter and English fan translation.
The intended use-case is for document translation tasks.
Dataset format:
{
'src': 'JAPANESE WEB NOVEL CHAPTER',
'trg': 'CORRESPONDING ENGLISH TRANSLATION',
'meta': {
'general': {
'series_title_eng': 'ENGLISH SERIES TITLE',
'series_title_jap': 'JAPANESE SERIES TITLE',
'sentence_alignment_score':… See the full description on the dataset page: https://huggingface.co/datasets/NilanE/ParallelFiction-Ja_En-100k.apr_rl_dataopen_parallel_think_cot_update_wo_answer
open_parallel_think_cot_update_wo_answer
This dataset is derived from haowu89/open_parallel_think_cot_update.
Transformation applied:
For every example, for every string item inside context, remove the final sentence.
The intent is to strip the trailing answer-bearing sentence while keeping the earlier reasoning trajectory.
Generated on 2026-04-15.
BFCL-V4-Parallel-Multi-Turn
BFCL V4 Parallel Multi-Turn
Flattened current-turn rows from BFCL v4 multi-turn trajectories for decentralized multi-agent function-calling experiments.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
turn_index
Categories
multi_turn_base_step
multi_turn_long_context_step
multi_turn_miss_func_step… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Multi-Turn.countdown_problemscantonese-chinese-parallel-corpus
Dataset Summary
This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation.
The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation.
Languages
Cantonese (yue)
Simplified Chinese (zh)
Dataset Structure
Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.dogv_parallel
DOGV_PARALLEL Dataset
Dataset Summary
DOGV_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research.
Dataset Structure
Each row in the dataset includes the following columns:
VA: A sentence in Valencian.
ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/dogv_parallel.parallel-translation-training-pool
Parallel translation training pool
Sentences in eleven languages beside their translations, from five public parallel corpora read at
the pinned revisions named below and laid out twice. Ten languages are paired with English in both
directions, twenty directions in all. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 4975238 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/parallel-translation-training-pool.ENG_NEP_MED_PARALLEL
Dataset Card for Dataset Name
This dataset aims to be a state of art Nepali English Parallel translation in Medical Domain. Further on this dataset will be updated with more correct translations.
Dataset Details
Dataset Description
Curated by: Bibek Poudel
Language(s) (NLP): Nepali, English
Dataset Sources [optional]
Repository: https://www.kaggle.com/datasets/rxnach/nepali-health-forum-corpus-questions-and-answers
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Bibek-Poudel/ENG_NEP_MED_PARALLEL.CAT-Parallel-Dataset-Indicsosp_sft_datamyanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
v3_msmarco_parallelai_e5qwen7b_6intent_claim_degrademulti-parallel-data
