CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mathoctopus /GSM8KInstruct_Paralleltextquestion-answering10K<n<100K11 likes1.3k downloads3y agoHugging Face02oss-codes /Cyber-Parallel-Dataset-Indictext1K<n<10K0 likes1k downloads1y agoHugging Face03oss-codes /Finance-Parallel-Dataset-Indictext100K<n<1M0 likes822 downloads1y agoHugging Face04Lego-MT /Parallel_Dataset Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: https://aclanthology.org/2025.findings-acl.1200.pdf Repository: https://github.com/CONE-MT/CONE texttranslation1M<n<10M0 likes755 downloads1y agoHugging Face05oss-codes /Law-Parallel-Dataset-Indictext100K<n<1M0 likes625 downloads1y agoHugging Face06Sudehsna /Romansh_German_Parallel_Data Romansh–German Parallel Dataset (FineWeb-Based) This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction. Description This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.tabular10K<n<100K2 likes472 downloads1y agoHugging Face07oss-codes /Medical-Parallel-Dataset-Indictext10K<n<100K0 likes327 downloads1y agoHugging Face08PumpkinCat /ParallelThinkingDLMtabular100K<n<1M0 likes319 downloads11mo agoHugging Face09oss-codes /CA-Parallel-Dataset-Indictext100K<n<1M0 likes241 downloads1y agoHugging Face10bashkorttele /trilingual-parallel-phrasebooks-bgpu Bashkir Trilingual Parallel Phrasebooks 9,857 phrases aligned across three languages — Bashkir, Russian and one of Altai, Arabic, Kazakh, Yakut (Sakha), Chinese — from five phrasebooks published by M. Akmulla Bashkir State Pedagogical University. One row is one phrase in all three languages: a parallel corpus for machine translation and cross-lingual work with a low-resource Turkic language. Each phrasebook is a separate file and a separate config, because the third language… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/trilingual-parallel-phrasebooks-bgpu.texttranslation1K<n<10K2 likes210 downloads15d agoHugging Face11raptorkwok /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M17 likes189 downloads3y agoHugging Face12OpenMLRL /BFCL-V4-Parallel-Native BFCL V4 Parallel Native Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration. Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files. Fields id official_category task_type user_prompt function ground_truth Categories live_parallel live_parallel_multiple parallel parallel_multiple Counts train: 352 rows eval: 88 rows total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.texttext-generationn<1K1 likes156 downloads3mo agoHugging Face13Nart /parallel_ab-ru Dataset Summary The Abkhaz Russian parallel corpus dataset is a collection of 205,665 sentences/words extracted from different sources; e-books, web scrapping. Dataset Creation Source Data Here is a link to the source on github Considerations for Using the Data Other Known Limitations The accuracy of the dataset is around 95% (gramatical, arthographical errors) texttext-generationn<1K1 likes134 downloads2y agoHugging Face14oss-codes /Jee-Parallel-Dataset-Indictext10K<n<100K0 likes128 downloads1y agoHugging Face15Mo-Abdalkader /Egyptian-Arabic-English-Parallel-Corpus Egyptian Arabic-English Parallel Corpus Author: Mohamed Abdalkader · LinkedIn · GitHub A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation. Dataset Structure egyptian-arabic-english-parallel-corpus/ ├── SFT/ │ ├── Train/ │ │ ├── topics/ # 1,800 individual topic JSON files │ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.text100K<n<1M1 likes123 downloads4d agoHugging Face16NilanE /ParallelFiction-Ja_En-100k Dataset details: Each entry in this dataset is a sentence-aligned Japanese web novel chapter and English fan translation. The intended use-case is for document translation tasks. Dataset format: { 'src': 'JAPANESE WEB NOVEL CHAPTER', 'trg': 'CORRESPONDING ENGLISH TRANSLATION', 'meta': { 'general': { 'series_title_eng': 'ENGLISH SERIES TITLE', 'series_title_jap': 'JAPANESE SERIES TITLE', 'sentence_alignment_score':… See the full description on the dataset page: https://huggingface.co/datasets/NilanE/ParallelFiction-Ja_En-100k.texttranslation100K<n<1M82 likes116 downloads2y agoHugging Face17Parallel-Reasoning /apr_rl_datatext100K<n<1M0 likes115 downloads1y agoHugging Face18haowu89 /open_parallel_think_cot_update_wo_answer open_parallel_think_cot_update_wo_answer This dataset is derived from haowu89/open_parallel_think_cot_update. Transformation applied: For every example, for every string item inside context, remove the final sentence. The intent is to strip the trailing answer-bearing sentence while keeping the earlier reasoning trajectory. Generated on 2026-04-15. text1K<n<10K0 likes104 downloads5mo agoHugging Face19OpenMLRL /BFCL-V4-Parallel-Multi-Turn BFCL V4 Parallel Multi-Turn Flattened current-turn rows from BFCL v4 multi-turn trajectories for decentralized multi-agent function-calling experiments. Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files. Fields id official_category task_type user_prompt function ground_truth turn_index Categories multi_turn_base_step multi_turn_long_context_step multi_turn_miss_func_step… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Multi-Turn.texttext-generation1K<n<10K0 likes104 downloads3mo agoHugging Face20Parallel-Reasoning /countdown_problemstabular100K<n<1M0 likes101 downloads1y agoHugging Face21HKAllen /cantonese-chinese-parallel-corpus Dataset Summary This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation. The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation. Languages Cantonese (yue) Simplified Chinese (zh) Dataset Structure Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.texttranslation100K<n<1M3 likes95 downloads2y agoHugging Face22gplsi /dogv_parallel DOGV_PARALLEL Dataset Dataset Summary DOGV_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research. Dataset Structure Each row in the dataset includes the following columns: VA: A sentence in Valencian. ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/dogv_parallel.texttranslation100K<n<1M1 likes85 downloads7mo agoHugging Face23Emulated-Inc /parallel-translation-training-pool Parallel translation training pool Sentences in eleven languages beside their translations, from five public parallel corpora read at the pinned revisions named below and laid out twice. Ten languages are paired with English in both directions, twenty directions in all. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 4975238 rows, one JSON object per line, with these fields. Field What it holds id a row identifier… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/parallel-translation-training-pool.texttranslation1M<n<10M0 likes84 downloads10d agoHugging Face24Bibek-Poudel /ENG_NEP_MED_PARALLEL Dataset Card for Dataset Name This dataset aims to be a state of art Nepali English Parallel translation in Medical Domain. Further on this dataset will be updated with more correct translations. Dataset Details Dataset Description Curated by: Bibek Poudel Language(s) (NLP): Nepali, English Dataset Sources [optional] Repository: https://www.kaggle.com/datasets/rxnach/nepali-health-forum-corpus-questions-and-answers Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Bibek-Poudel/ENG_NEP_MED_PARALLEL.text10K<n<100K1 likes82 downloads1y agoHugging Face25oss-codes /CAT-Parallel-Dataset-Indictext10K<n<100K0 likes80 downloads1y agoHugging Face26Parallel-Reasoning /sosp_sft_datatabular100K<n<1M0 likes78 downloads1y agoHugging Face27freococo /myanmar_quran_parallel_dataset_human_vs_ai Myanmar Quran Parallel Dataset: Human vs AI This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses. It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context. Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.texttranslation1K<n<10K0 likes70 downloads8mo agoHugging Face28ICML-2026-agent-repro /repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes70 downloads2mo agoHugging Face29Nithish2410 /v3_msmarco_parallelai_e5qwen7b_6intent_claim_degradetext100K<n<1M0 likes70 downloads1mo agoHugging Face30adhakimi /multi-parallel-datatext100M<n<1B0 likes63 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.