CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlidarAsvarov /dagestan-constitution Constitution of the Republic of Dagestan — 13 languages Scans of the Constitution of the Republic of Dagestan in thirteen languages: eleven indigenous languages of Dagestan and the North Caucasus, plus Azerbaijani and Russian. 291 PDF pages across 13 files. Most PDF pages are two-page book spreads, not single pages. 245 of the 291 are landscape scans of an open book, so the corpus is really 536 book pages. Anyone building an OCR pipeline needs to split them — see Per-file… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/dagestan-constitution.tabularimage-to-textn<1K1 likes520 downloads2mo agoHugging Face02sequelbox /DAG-Reasoning-DeepSeek-R1-0528Click here to support our open-source dataset and model releases! DAG-Reasoning-DeepSeek-R1-0528 is a dataset focused on analysis and reasoning, creating directed acyclic graphs testing the limits of DeepSeek R1 0528's graph-reasoning skills! This dataset contains: 4.08k synthetically generated prompts to create directed acyclic graphs in response to user input, with all responses generated using DeepSeek R1 0528. All responses contain a multi-step thinking process to perform effective… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DAG-Reasoning-DeepSeek-R1-0528.texttext-generation1K<n<10K12 likes76 downloads1y agoHugging Face03xri /dagaare_synth_trainThe dagaare_dict_guided_train_9k.tsv presents the synthesized data for the dictionary-guided training set, which is designed on the observed dictionary "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021). Dictionary terms were selected as target words in the curation of the dataset. text1K<n<10K0 likes19 downloads1y agoHugging Face04abdulhafis /Dagbani_English_Datasettext1K<n<10K0 likes19 downloads6mo agoHugging Face05xri /dagaare_synth_source_targetThe dagaareDictTrain.tsv was generated using the Machine Translation from One Book (MTOB) technique using "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021) for LLM context. tabular1K<n<10K0 likes15 downloads1y agoHugging Face06dagn /expanded-amharic-news-datasetgated Expanded Amharic News Dataset (2011–2024) Dataset Description The Expanded Amharic News Dataset is a large-scale, ethically collected corpus of Amharic-language news articles written in Geʽez (Fidel) script, designed to support research in Natural Language Processing (NLP). This dataset builds upon the “An Amharic News Text Classification Dataset” developed by Israel Abebe Azime and Nebil Mohammed (arXiv link), which categorized Amharic news articles into multiple topical… See the full description on the dataset page: https://huggingface.co/datasets/dagn/expanded-amharic-news-dataset.texttext-classification100K<n<1M0 likes12 downloads9mo agoHugging Face07dagmawi-ml /cwe-workshop-datasetimage1M<n<10M0 likes10 downloads4mo agoHugging Face08Talelaw /dagpengertextn<1K0 likes8 downloads3y agoHugging Face09AlidarAsvarov /dagwoman-parallelgated Женщина Дагестана — Russian–Dagestani parallel corpus Sentence pairs mined from the print archive of Женщина Дагестана ("Woman of Dagestan"), a magazine published in Russian and six Cyrillic-script Dagestani languages. 21,615 rows, issues 2016–2026. One row per Russian sentence; each language's rendering of it sits in its own column, empty where that edition does not carry it. column language sentences rus Russian (pivot) 21,615 lez Lezgian 21,615 Three columns… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/dagwoman-parallel.texttranslation10K<n<100K0 likes3 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.