datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code2doc
Code2Doc: Function-Documentation Pairs Dataset
A curated dataset of 13,358 high-quality function-documentation pairs extracted from popular open-source repositories on GitHub. Designed for training models to generate documentation from code.
Dataset Description
This dataset contains functions paired with their docstrings/documentation comments from 5 programming languages, extracted from well-maintained, highly-starred GitHub repositories.
Languages Distribution… See the full description on the dataset page: https://huggingface.co/datasets/kaanrkaraman/code2doc.kknews-dataset
KKNews.uz Dataset
Qaraqalpaqstan Xabar Agentligi (kknews.uz) maqalaları — 5 tilde.
Languages
Code
Language
ru
Russian
uz
Uzbek (Latin)
oz
Uzbek (Cyrillic)
kk
Karakalpak (Cyrillic)
qq
Karakalpak (Latin)
Columns
Column
Type
Description
id
int
WordPress post ID
lang
string
Language code
category_id
int
Category ID
category_name
string
Category name
title
string
Plain text title
content_html
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/kaa-ml/kknews-dataset.kaa-parallel-corpus
Kaa Karakalpak-English Parallel Corpus (FineTranslations)
📌 Overview
This repository contains a high-quality, curated parallel corpus for the Karakalpak (kaa) language, paired with English (en). Karakalpak is a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan.
This dataset is a specialized subset extracted from the massive HuggingFaceFW/finetranslations project. The goal of this repo is to provide a dedicated and easy-to-access resource… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/kaa-parallel-corpus.
