datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dagestan-constitution
Constitution of the Republic of Dagestan — 13 languages
Scans of the Constitution of the Republic of Dagestan in thirteen languages: eleven
indigenous languages of Dagestan and the North Caucasus, plus Azerbaijani and Russian.
291 PDF pages across 13 files.
Most PDF pages are two-page book spreads, not single pages. 245 of the 291 are
landscape scans of an open book, so the corpus is really 536 book pages. Anyone
building an OCR pipeline needs to split them — see Per-file… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/dagestan-constitution.DAG-Reasoning-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
DAG-Reasoning-DeepSeek-R1-0528 is a dataset focused on analysis and reasoning, creating directed acyclic graphs testing the limits of DeepSeek R1 0528's graph-reasoning skills!
This dataset contains:
4.08k synthetically generated prompts to create directed acyclic graphs in response to user input, with all responses generated using DeepSeek R1 0528.
All responses contain a multi-step thinking process to perform effective… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DAG-Reasoning-DeepSeek-R1-0528.dagaare_synth_trainThe dagaare_dict_guided_train_9k.tsv presents the synthesized data for the dictionary-guided training set, which is designed on the observed dictionary "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021).
Dictionary terms were selected as target words in the curation of the dataset.
Dagbani_English_Datasetdagaare_synth_source_targetThe dagaareDictTrain.tsv was generated using the Machine Translation from One Book (MTOB) technique using "A dictionary and grammatical sketch of Dagaare" by Ali, Grimm, and Bodomo (2021) for LLM context.
expanded-amharic-news-dataset
Expanded Amharic News Dataset (2011–2024)
Dataset Description
The Expanded Amharic News Dataset is a large-scale, ethically collected corpus of Amharic-language news articles written in Geʽez (Fidel) script, designed to support research in Natural Language Processing (NLP).
This dataset builds upon the “An Amharic News Text Classification Dataset” developed by Israel Abebe Azime and Nebil Mohammed (arXiv link), which categorized Amharic news articles into multiple topical… See the full description on the dataset page: https://huggingface.co/datasets/dagn/expanded-amharic-news-dataset.cwe-workshop-datasetdagpengerdagwoman-parallel
Женщина Дагестана — Russian–Dagestani parallel corpus
Sentence pairs mined from the print archive of Женщина Дагестана ("Woman of Dagestan"), a
magazine published in Russian and six Cyrillic-script Dagestani languages.
21,615 rows, issues 2016–2026. One row per Russian sentence; each
language's rendering of it sits in its own column, empty where that edition does not carry it.
column
language
sentences
rus
Russian (pivot)
21,615
lez
Lezgian
21,615
Three columns… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/dagwoman-parallel.
