datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Orthography-and-Spelling
🇰🇿 Kazakh Orthographic and Quantifier Refinement
📖 Overview
This dataset contains 1,504 samples focusing on common orthographic and grammatical nuances in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
1,504
Total Words (approx.)
68,469
Avg. Words per Sample
45
Word Count Distribution (Per Field)
The following table details the distribution of word counts… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Orthography-and-Spelling.AAAI_Swahili_dataset
README for Swahili Translated Dataset from Toloka
Dataset Description
This dataset is a Dolly 15k translated from English to Swahili, filtered and processed using the Toloka platform. It includes various contexts, responses, and instructions from diverse domains, providing a rich resource for natural language processing tasks, particularly for those focusing on the Swahili language.
Data Fields
task_id: A unique identifier for each task in the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/ortofasfat/AAAI_Swahili_dataset.
