datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kazakh-news-summarization-20k-adapted
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
kazakh_news_summarization
This dataset contains pairs of Kazakh language prompts and completions focused on summarizing news articles from sources like BAQ.KZ. The content covers diverse topics including social issues, legal cases, government initiatives, and international events within Kazakhstan and abroad. Each entry consists of a standard instruction to summarize text… See the full description on the dataset page: https://huggingface.co/datasets/shayekh/kazakh-news-summarization-20k-adapted.resume-summarization-dataset
Resume Summarization Dataset
This dataset contains machine-generated summaries of 14,505 resumes using gpt-4o-mini. Each entry includes the original resume and a markdown-formatted summary divided into 5 sections.
Structure
Each row is a JSON object with:
resume: The original resume text
summary: The structured markdown summary
input_tokens and output_tokens: (optional) token usage info
License
Some portions of this dataset are derived from public sources… See the full description on the dataset page: https://huggingface.co/datasets/jbeiroa/resume-summarization-dataset.code-text-galeras-code-summarization-3k-dedupedrussian_summarizationplora-clean-summarization
PLoRA Clean Multilingual Summarization
This dataset is a cached, length-filtered training bundle for the local PLoRA
notebook. It contains prompt/answer records whose full Qwen chat-formatted
sequence length is at most 4096 tokens.
Languages: hin_Deva (Hindi), fra_Latn (French), cmn_Hans (Chinese), urd_Arab (Urdu), eng_Latn (English), nld_Latn (Dutch), pol_Latn (Polish), snd_Arab (Sindhi), ben_Beng (Bengali), mar_Deva (Marathi)
Counts:
Train records: 100,000
Validation records: 8… See the full description on the dataset page: https://huggingface.co/datasets/divyajot5005/plora-clean-summarization.
