CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K1 likes188 downloads2y agoHugging Face02abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes57 downloads3mo agoHugging Face03DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes47 downloads4mo agoHugging Face04ClarusC64 /clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1 Goal Test if a model can hold separate reasoning streams at once Detect constraint dismissal Detect bleed-over where one stream turns into claims in the other What it measures streams_heldResponse acknowledges and maintains both streams bleed_overConstraint stream improperly becomes a medical claim, or vice versa premature_synthesisResponse forces a single solution that silences one stream assumption_collapseResponse drops a premise entirely Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.texttext-generationn<1K0 likes42 downloads8mo agoHugging Face05abdelhaqueidali /Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset Dataset Card for Tachelhit Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Tachelhit language ($\text{Tacelḥit}$ / $\text{Tamazigt}$), pairing native Latin-based orthography with automated Amazigh script transliterations. It is built by processing clean source sentences through an algorithmic engine designed to handle phonetic mappings, manage contextual schwa distributions, and safeguard acronyms and foreign proper names. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset.texttranslation10K<n<100K0 likes36 downloads3mo agoHugging Face06ReliableAI /Irish-English-Parallel-Collection UCCIX's English-Irish Parallel Textual Corpus Dataset Summary This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR. This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data. Dataset Sources Source Description Statistics Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.texttext-generation10K<n<100K1 likes32 downloads2y agoHugging Face07bekan /english_karakalpak_parallel_corpus_v1 English-Karakalpak Parallel Corpus (en-kaa) Dataset Description English-Karakalpak Parallel Corpus is a high-quality dataset containing 10,441 aligned sentence pairs in English and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. The corpus utilizes the official… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v1.texttranslation10K<n<100K2 likes12 downloads10mo agoHugging Face08bekan /english_karakalpak_pairs_parallel_corpus_v2_8907 English-Karakalpak Parallel Corpus v2 (8.9K) Dataset Description English-Karakalpak Parallel Corpus v2 is a high-quality dataset containing 8,906 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. This resource… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_pairs_parallel_corpus_v2_8907.texttranslation1K<n<10K1 likes4 downloads10mo agoHugging Face09Dddixyy /Italian_latin_parallel_animals descrizioni di animali e habitat - Synthetic Dataset This dataset was generated using the Synthetic Dataset Generator powered by Gemini AI. Topic: descrizioni di animali e habitat Field 1: italiano Field 2: latino antico(traduzione) Rows: 280 Generated on: 2025-05-27T00:07:49.042Z texttext-generationn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.