datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ParallelKernelBench_Problems
ParallelKernelBench (benchmark)
Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels.
This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Inputs are deterministic — reproduce them with create_input_tensor(rank, world_size, problem_id, base_shape, dtype, trial) from that file; you do not need stored .pt files.
Files
Path
Description… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/ParallelKernelBench_Problems.ParallelKernelBench_Problems
ParallelKernelBench (benchmark)
Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels.
This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py.
Files
Path
Description
data/problems.parquet
One row per problem (tabular access)
reference/*.py
Reference solution() implementations
utils/input_output_tensors.py
Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.scipar_parallel_docs
SciPar Parallel Documents
Dataset Description
This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts.
In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories.
This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences.
To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tunisian-msa-parallel-corpus
Dataset Description
This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models.
The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.ParallelKernelBench_Kernels
ParallelKernelBench Kernels
Net-new multi-GPU CUDA kernels generated by LLMs for ParallelKernelBench.
Each subdirectory under solutions/ is one model run. File names match the benchmark problem stems (e.g. 17_rope_allgather_cuda.py ↔ problem 17_rope_allgather in willychan21/ParallelKernelBench_Problems).
Layout
solutions/
<run_id>/
<stem>_cuda.py
...
Runs (1 run(s), 87 kernel files)
run_id
kernels
path… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Kernels.quran-parallel-corpus
Quran Parallel Corpus
Verse-aligned Quran parallel corpus — Arabic (Uthmani), English (Sahih International), and Indonesian (Ministry of Religious Affairs).
Stats
Total verses: 6236
Languages: Arabic, English, Indonesian
Translation pairs: Arabic↔English, Arabic↔Indonesian, English↔Indonesian
Formats: JSONL, CSV, Parquet
Structure
Each verse record contains:
Field
Description
surah_number
Chapter (1–114)
surah_name_arabic
Arabic surah… See the full description on the dataset page: https://huggingface.co/datasets/sarjukesumo/quran-parallel-corpus.human-ai-parallel-detection
Dataset Card for human-ai-parallel-detection
Dataset Description
Dataset Summary
The human-ai-parallel-detection dataset contains 600 balanced instances for evaluating methods to distinguish between human-written and AI-generated text continuations. Each instance includes a 500-word human-written prompt followed by parallel continuations from humans, GPT-4o, and LLaMA-70B-Instruct. The dataset includes both style embedding features and LLM-as-judge predictions… See the full description on the dataset page: https://huggingface.co/datasets/ephipi/human-ai-parallel-detection.bible-parallel-english
Parallel Bible — English Translations and Ancient Versions
A verse-aligned parallel corpus of the Protestant Bible in seventeen English
translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta
for the New Testament.
Looking for every language? This repository is a curated English set,
chosen for spread across translation families and small enough to load whole.
For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses —
see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.kaa-parallel-corpus
Kaa Karakalpak-English Parallel Corpus (FineTranslations)
📌 Overview
This repository contains a high-quality, curated parallel corpus for the Karakalpak (kaa) language, paired with English (en). Karakalpak is a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan.
This dataset is a specialized subset extracted from the massive HuggingFaceFW/finetranslations project. The goal of this repo is to provide a dedicated and easy-to-access resource… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/kaa-parallel-corpus.tunisian-msa-parallel-corpus-evaluated
Dataset Description
This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb).
It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP.
The primary goals are to support:
Machine translation between Tunisian Arabic and MSA.
Research in dialectal-aware text generation and evaluation.
Cross-dialect representation learning in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus-evaluated.bashkir-russian-parallel
Dataset Card for Bashkir-Russian Parallel Corpus
Dataset Details
Dataset Description
Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart.
The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.ParallelPrompt
PARALLELPROMPT
A benchmark dataset of 37,021 parallelizable prompts from real-world LLM conversations, designed for optimizing LLM serving systems through intra-query parallelism.
Repository and Resources
Dataset: Hugging Face
Code: GitHub
Paper: PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries
The GitHub repository contains:
Data curation pipeline
Schema extraction code
Evaluation suite for measuring latency and quality
Baseline implementations… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/ParallelPrompt.parason-data
parason-data
Evaluation traces and structural-analysis artifacts for the parallel-reasoning line of work.
Training data is not here — it lives in parallel-reasoner/sft-ours (splits 1x, 8x).
This repo holds generated traces, so that structural claims about model behaviour can be
re-derived rather than taken on trust.
Layout
aime24/<model-name>/traces.jsonl
aime24/Qwen3-8B-sft-ours8x-ar/
Traces from parallel-reasoner/Qwen3-8B-sft-ours8x-ar
— the… See the full description on the dataset page: https://huggingface.co/datasets/parallel-reasoner/parason-data.dhivehi-legal-text-parallel
Dhivehi-English Legal Parallel Corpus
Dataset Description
A high-quality parallel corpus of 56,556 Dhivehi-English sentence pairs extracted from 200 Maldivian legal documents. This dataset is deduplicated and cleaned for machine translation and bilingual model training.
Dataset Summary
Languages: Dhivehi (dv) ↔ English (en)
Total Pairs: 56,556
Source Laws: 200
Duplicates Removed: 31,235
Average Dhivehi Length: 173.6 characters
Average English… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-legal-text-parallel.tunisian-msa-parallel-corpus-evaluated
Dataset Description
This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb).
It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP.
The primary goals are to support:
Machine translation between Tunisian Arabic and MSA.
Research in dialectal-aware text generation and evaluation.
Cross-dialect representation learning in Arabic… See the full description on the dataset page: https://huggingface.co/datasets/hbenayed/tunisian-msa-parallel-corpus-evaluated.parallel-bm-en
khursanirevo/parallel-bm-en
Parallel English-Bahasa Melayu translation pairs (102k rows, OpenHermes-derived).
Splits
split
rows
train
97,280
validation
5,120
Stratified 95/5 by source/category (seed=42).
Source files
data/sft/parallel_bm_en_30m.jsonl
Schema
Each row is a JSON object. See the loader script for field details.
Provenance
Generated as part of MaLLaM 2026 Tiny pretraining/SFT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/parallel-bm-en.
