CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Sudehsna /Romansh_German_Parallel_Data Romansh–German Parallel Dataset (FineWeb-Based) This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction. Description This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.tabular10K<n<100K2 likes469 downloads1y agoHugging Face02PumpkinCat /ParallelThinkingDLMtabular100K<n<1M0 likes319 downloads11mo agoHugging Face03Parallel-Reasoning /countdown_problemstabular100K<n<1M0 likes103 downloads1y agoHugging Face04Parallel-Reasoning /sosp_sft_datatabular100K<n<1M0 likes78 downloads1y agoHugging Face05ICML-2026-agent-repro /repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes70 downloads2mo agoHugging Face06sermonindex /bible-parallel-english Parallel Bible — English Translations and Ancient Versions A verse-aligned parallel corpus of the Protestant Bible in seventeen English translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta for the New Testament. Looking for every language? This repository is a curated English set, chosen for spread across translation families and small enough to load whole. For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses — see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.tabulartext-generation100K<n<1M0 likes49 downloads12d agoHugging Face07Shiki42 /screw_retimed_parallel_23_no_fastforward_20260807tabularn<1K0 likes41 downloads2mo agoHugging Face08alaminerca /sango-french-bible-parallel SFPC: Sango-French Parallel Corpus The first quality-filtered, verse-aligned Sango-French parallel corpus, constructed for neural machine translation research. This dataset directly addresses the "Sango Problem" identified by Meta's NLLB-200 project — the failure of cross-lingual transfer for a linguistically isolated Creole language. Associated resources: Model: alaminerca/nllb-sango-french Demo: Sango-French Translator Paper: SangoNMT: Parameter-Efficient Domain Adaptation of… See the full description on the dataset page: https://huggingface.co/datasets/alaminerca/sango-french-bible-parallel.tabulartranslation10K<n<100K1 likes34 downloads5mo agoHugging Face09BashkirNLPWorld /bashkir-russian-parallelgated Dataset Card for Bashkir-Russian Parallel Corpus Dataset Details Dataset Description Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart. The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.tabulartext-generation1M<n<10M0 likes27 downloads5d agoHugging Face10freococo /vinaya-pitaka-pali-myanmar-parallel Vinaya Pitaka: Pali-Myanmar Parallel Dataset Description This dataset provides a professionally aligned, paragraph-level parallel corpus of the Vinaya Pitaka (The Code of Monastic Discipline). It features the original Pali text (presented in Myanmar script) alongside its modern Myanmar translation. The dataset covers all five major volumes of the Vinaya: Pārājika (ပါရာဇိကပါဠိ / ပါရာဇိကဏ်) Pācittiya (ပါစိတ္တိယပါဠိ / ပါစိတ်) Mahāvagga (မဟာဝဂ္ဂပါဠိ / မဟာဝါ) Cūḷavagga… See the full description on the dataset page: https://huggingface.co/datasets/freococo/vinaya-pitaka-pali-myanmar-parallel.tabulartranslation1K<n<10K0 likes23 downloads8mo agoHugging Face11enesyila /ota-bible-parallel Ottoman–Turkish–English Parallel New Testament Corpus This repository is a verse-aligned parallel corpus for Ottoman Turkish (Perso-Arabic original script) with English and modern Turkish reference translations. It is intended for training and evaluating translation models for Ottoman Turkish, a low-resource historical language. Source language: Ottoman Turkish (ota), Perso-Arabic script Target languages: English (en), modern Turkish (tr) Unit of alignment: a single New… See the full description on the dataset page: https://huggingface.co/datasets/enesyila/ota-bible-parallel.tabulartranslationn<1K0 likes21 downloads1mo agoHugging Face12parallel-reasoner /parason-data parason-data Evaluation traces and structural-analysis artifacts for the parallel-reasoning line of work. Training data is not here — it lives in parallel-reasoner/sft-ours (splits 1x, 8x). This repo holds generated traces, so that structural claims about model behaviour can be re-derived rather than taken on trust. Layout aime24/<model-name>/traces.jsonl aime24/Qwen3-8B-sft-ours8x-ar/ Traces from parallel-reasoner/Qwen3-8B-sft-ours8x-ar — the… See the full description on the dataset page: https://huggingface.co/datasets/parallel-reasoner/parason-data.tabulartext-generationn<1K0 likes17 downloads2mo agoHugging Face13ssergaroo /english-classics-parallel-samples Booklern English classics: parallel samples Paragraph-aligned opening passages of public-domain English classics with a translation into Spanish, Japanese, Brazilian Portuguese, Russian, Chinese, published by Booklern, a bilingual book reader for learning English through real books. Each book is read on Booklern with a sentence-by-sentence translation under the English, read-aloud audio, a dictionary and vocabulary tools; the rows here are the same opening paragraphs that appear… See the full description on the dataset page: https://huggingface.co/datasets/ssergaroo/english-classics-parallel-samples.tabulartranslation1K<n<10K0 likes13 downloads14h agoHugging Face14siberian-lang-lab /evenki-rus-parallel-corporatabular1K<n<10K0 likes12 downloads10mo agoHugging Face15narinzar /parallel-image-text-dataset-builder parallel-image-text-dataset-builder (sample) A small representative sample from the parallel-image-text-dataset-builder pipeline: it ingests image-text pairs, removes near-duplicates with perceptual-hash (dhash) LSH-style bucketing, filters weak pairs by CLIP image-text similarity, and writes fixed-size WebDataset-style tar shards. Contents shard-00002.tar - one WebDataset-style shard (536 samples). Each sample is two members sharing a key: {key}.jpg (image) and… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/parallel-image-text-dataset-builder.tabularimage-to-textn<1K0 likes11 downloads3mo agoHugging Face16khursanirevo /parallel-bm-en khursanirevo/parallel-bm-en Parallel English-Bahasa Melayu translation pairs (102k rows, OpenHermes-derived). Splits split rows train 97,280 validation 5,120 Stratified 95/5 by source/category (seed=42). Source files data/sft/parallel_bm_en_30m.jsonl Schema Each row is a JSON object. See the loader script for field details. Provenance Generated as part of MaLLaM 2026 Tiny pretraining/SFT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/parallel-bm-en.tabulartext-generation100K<n<1M0 likes10 downloads3mo agoHugging Face17abhinandansamal /odia-german-parallel-corpus-researchgated Dataset Summary This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics. The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.tabulartranslation1K<n<10K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.