datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kazakh-news-summarization-20k-adapted
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
kazakh_news_summarization
This dataset contains pairs of Kazakh language prompts and completions focused on summarizing news articles from sources like BAQ.KZ. The content covers diverse topics including social issues, legal cases, government initiatives, and international events within Kazakhstan and abroad. Each entry consists of a standard instruction to summarize text… See the full description on the dataset page: https://huggingface.co/datasets/shayekh/kazakh-news-summarization-20k-adapted.ainavox-kazakh-preprocessed
AinaVox Kazakh TTS preprocessed training artifacts
Precomputed training artifacts used for the AinaVox Kazakh IndexTTS-2
experiments. This repository is intended to avoid repeating the expensive
feature-extraction stage when reproducing or extending the training runs.
The binary artifacts are stored in the public Hugging Face Storage Bucket
ruslawik/ainavox-kazakh-preprocessed-data.
This dataset repository contains the documentation and source integrity
manifest.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ruslawik/ainavox-kazakh-preprocessed.
