CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01togethercomputer /ParallelKernelBench_Problems ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Inputs are deterministic — reproduce them with create_input_tensor(rank, world_size, problem_id, base_shape, dtype, trial) from that file; you do not need stored .pt files. Files Path Description… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/ParallelKernelBench_Problems.tabulartext-generationn<1K0 likes362 downloads3mo agoHugging Face02adeshkin /khakas-russian-parallel-corpus Khakas-Russian Parallel Corpus The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people. Dataset Overlap: The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.texttranslation100K<n<1M2 likes277 downloads12d agoHugging Face03willychan21 /ParallelKernelBench_Problems ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.tabulartext-generationn<1K0 likes266 downloads4mo agoHugging Face04ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K1 likes188 downloads2y agoHugging Face05OpenMLRL /BFCL-V4-Parallel-Native BFCL V4 Parallel Native Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration. Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files. Fields id official_category task_type user_prompt function ground_truth Categories live_parallel live_parallel_multiple parallel parallel_multiple Counts train: 352 rows eval: 88 rows total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.texttext-generationn<1K1 likes156 downloads3mo agoHugging Face06Nart /parallel_ab-ru Dataset Summary The Abkhaz Russian parallel corpus dataset is a collection of 205,665 sentences/words extracted from different sources; e-books, web scrapping. Dataset Creation Source Data Here is a link to the source on github Considerations for Using the Data Other Known Limitations The accuracy of the dataset is around 95% (gramatical, arthographical errors) texttext-generationn<1K1 likes134 downloads2y agoHugging Face07OpenMLRL /BFCL-V4-Parallel-Multi-Turn BFCL V4 Parallel Multi-Turn Flattened current-turn rows from BFCL v4 multi-turn trajectories for decentralized multi-agent function-calling experiments. Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files. Fields id official_category task_type user_prompt function ground_truth turn_index Categories multi_turn_base_step multi_turn_long_context_step multi_turn_miss_func_step… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Multi-Turn.texttext-generation1K<n<10K0 likes104 downloads3mo agoHugging Face08tunis-ai /tunisian-msa-parallel-corpus Dataset Description This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models. The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.tabulartranslation1K<n<10K0 likes98 downloads1y agoHugging Face09SINAI /ALIA-parallel-translation Dataset Card for ALIA Parallel Translation Corpus This corpus comprises 35,753,765 domain-specific parallel segments (Spanish-English) designed for training and evaluating machine translation models in specialized domains. The corpus includes three main domains: Legal-Administrative, Biomedical, and Heritage, carefully curated to support document-level and multi-paragraph translation tasks beyond traditional sentence-level approaches. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-parallel-translation.texttranslation10M<n<100M0 likes93 downloads5mo agoHugging Face10Emulated-Inc /parallel-translation-training-pool Parallel translation training pool Sentences in eleven languages beside their translations, from five public parallel corpora read at the pinned revisions named below and laid out twice. Ten languages are paired with English in both directions, twenty directions in all. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 4975238 rows, one JSON object per line, with these fields. Field What it holds id a row identifier… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/parallel-translation-training-pool.texttranslation1M<n<10M0 likes84 downloads10d agoHugging Face11freococo /myanmar_quran_parallel_dataset_human_vs_ai Myanmar Quran Parallel Dataset: Human vs AI This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses. It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context. Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.texttranslation1K<n<10K0 likes70 downloads8mo agoHugging Face12willychan21 /ParallelKernelBench_Kernels ParallelKernelBench Kernels Net-new multi-GPU CUDA kernels generated by LLMs for ParallelKernelBench. Each subdirectory under solutions/ is one model run. File names match the benchmark problem stems (e.g. 17_rope_allgather_cuda.py ↔ problem 17_rope_allgather in willychan21/ParallelKernelBench_Problems). Layout solutions/ <run_id>/ <stem>_cuda.py ... Runs (1 run(s), 87 kernel files) run_id kernels path… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Kernels.tabulartext-generationn<1K0 likes68 downloads4mo agoHugging Face13avewright /exp085-parallel-multipv-harvest exp085 Parallel MultiPV Harvest This dataset is a frozen export of the exp085_parallel_multipv_harvest.py data harvester. Snapshot Export date: 2026-03-31 Shards: 44 Records: 224191 Size: 551330268 bytes of JSONL shard data Final shard: positions_000044.jsonl Final shard records: 2724 Included Files dataset/positions_*.jsonl manifest.json status.json exp085.log stdout.log seen_positions.sqlite Notes The JSONL line count is treated as the… See the full description on the dataset page: https://huggingface.co/datasets/avewright/exp085-parallel-multipv-harvest.text-generation100K<n<1M0 likes66 downloads6mo agoHugging Face14Malikeh1375 /complex_mathematical-scientific-notation-parallel Mathematical and Scientific Notation Parallel Corpus Dataset Description This dataset is designed for tokenizer robustness testing in mathematical and scientific contexts. It contains identical mathematical content expressed in four different notation styles, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison: Compare how different tokenizers (BPE, SentencePiece… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/complex_mathematical-scientific-notation-parallel.texttext-generationn<1K1 likes59 downloads1y agoHugging Face15Malikeh1375 /basic_mathematical-scientific-notation-parallel Mathematical and Scientific Notation Parallel Corpus Dataset Description This dataset is designed for tokenizer robustness testing in mathematical and scientific contexts. It contains identical mathematical content expressed in four different notation styles, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison: Compare how different tokenizers (BPE, SentencePiece… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/basic_mathematical-scientific-notation-parallel.texttext-generationn<1K0 likes57 downloads1y agoHugging Face16abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes57 downloads3mo agoHugging Face17sarjukesumo /quran-parallel-corpus Quran Parallel Corpus Verse-aligned Quran parallel corpus — Arabic (Uthmani), English (Sahih International), and Indonesian (Ministry of Religious Affairs). Stats Total verses: 6236 Languages: Arabic, English, Indonesian Translation pairs: Arabic↔English, Arabic↔Indonesian, English↔Indonesian Formats: JSONL, CSV, Parquet Structure Each verse record contains: Field Description surah_number Chapter (1–114) surah_name_arabic Arabic surah… See the full description on the dataset page: https://huggingface.co/datasets/sarjukesumo/quran-parallel-corpus.tabulartranslation100K<n<1M0 likes57 downloads1mo agoHugging Face18ephipi /human-ai-parallel-detection Dataset Card for human-ai-parallel-detection Dataset Description Dataset Summary The human-ai-parallel-detection dataset contains 600 balanced instances for evaluating methods to distinguish between human-written and AI-generated text continuations. Each instance includes a 500-word human-written prompt followed by parallel continuations from humans, GPT-4o, and LLaMA-70B-Instruct. The dataset includes both style embedding features and LLM-as-judge predictions… See the full description on the dataset page: https://huggingface.co/datasets/ephipi/human-ai-parallel-detection.tabulartext-classificationn<1K1 likes51 downloads1y agoHugging Face19sermonindex /bible-parallel-english Parallel Bible — English Translations and Ancient Versions A verse-aligned parallel corpus of the Protestant Bible in seventeen English translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta for the New Testament. Looking for every language? This repository is a curated English set, chosen for spread across translation families and small enough to load whole. For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses — see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.tabulartext-generation100K<n<1M0 likes49 downloads12d agoHugging Face20gplsi /alia_multilingual_parallel_sentences MULTILINGUAL PARALLEL SENTENCES Dataset The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models. It provides aligned sentences in multiple languages to facilitate multilingual learning. Dataset Structure The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language. The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.texttext-generation1M<n<10M0 likes48 downloads8mo agoHugging Face21DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes47 downloads4mo agoHugging Face22ClarusC64 /clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1 Goal Test if a model can hold separate reasoning streams at once Detect constraint dismissal Detect bleed-over where one stream turns into claims in the other What it measures streams_heldResponse acknowledges and maintains both streams bleed_overConstraint stream improperly becomes a medical claim, or vice versa premature_synthesisResponse forces a single solution that silences one stream assumption_collapseResponse drops a premise entirely Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.texttext-generationn<1K0 likes42 downloads8mo agoHugging Face23nickoo004 /kaa-parallel-corpus Kaa Karakalpak-English Parallel Corpus (FineTranslations) 📌 Overview This repository contains a high-quality, curated parallel corpus for the Karakalpak (kaa) language, paired with English (en). Karakalpak is a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan. This dataset is a specialized subset extracted from the massive HuggingFaceFW/finetranslations project. The goal of this repo is to provide a dedicated and easy-to-access resource… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/kaa-parallel-corpus.tabulartranslation10K<n<100K0 likes41 downloads5mo agoHugging Face24Lots-of-LoRAs /task1435_ro_sts_parallel_language_translation_ro_to_en Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1435_ro_sts_parallel_language_translation_ro_to_en Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1435_ro_sts_parallel_language_translation_ro_to_en.texttext-generation1K<n<10K0 likes39 downloads2y agoHugging Face25Lots-of-LoRAs /task1436_ro_sts_parallel_language_translation_en_to_ro Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1436_ro_sts_parallel_language_translation_en_to_ro Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1436_ro_sts_parallel_language_translation_en_to_ro.texttext-generation1K<n<10K0 likes38 downloads2y agoHugging Face26haowu89 /open_parallel_think_source Open Parallel Think — Source (per-model subsets) Math reasoning traces distilled from a shared question set by four models, organized one subset (config) per source model. Each question carries multiple reasoning traces ("parallel think"); here those traces are partitioned by the model that produced them. The underlying questions come from three collections: openmathinstruct, numinamath, and deepscale (the source is the prefix of guid, e.g. deepscale_10003). Subsets… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_source.texttext-generation100K<n<1M0 likes37 downloads3mo agoHugging Face27abdelhaqueidali /Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset Dataset Card for Tachelhit Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Tachelhit language ($\text{Tacelḥit}$ / $\text{Tamazigt}$), pairing native Latin-based orthography with automated Amazigh script transliterations. It is built by processing clean source sentences through an algorithmic engine designed to handle phonetic mappings, manage contextual schwa distributions, and safeguard acronyms and foreign proper names. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset.texttranslation10K<n<100K0 likes36 downloads3mo agoHugging Face28Lots-of-LoRAs /task441_eng_guj_parallel_corpus_gu-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task441_eng_guj_parallel_corpus_gu-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task441_eng_guj_parallel_corpus_gu-en_language_identification.texttext-generation1K<n<10K0 likes34 downloads2y agoHugging Face29tunis-ai /tunisian-msa-parallel-corpus-evaluated Dataset Description This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation. Cross-dialect representation learning in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus-evaluated.tabulartranslation1K<n<10K2 likes34 downloads1y agoHugging Face30ReliableAI /Irish-English-Parallel-Collection UCCIX's English-Irish Parallel Textual Corpus Dataset Summary This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR. This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data. Dataset Sources Source Description Statistics Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.texttext-generation10K<n<100K1 likes32 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.