CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shayekh /kazakh-news-summarization-20k-adapted This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. kazakh_news_summarization This dataset contains pairs of Kazakh language prompts and completions focused on summarizing news articles from sources like BAQ.KZ. The content covers diverse topics including social issues, legal cases, government initiatives, and international events within Kazakhstan and abroad. Each entry consists of a standard instruction to summarize text… See the full description on the dataset page: https://huggingface.co/datasets/shayekh/kazakh-news-summarization-20k-adapted.tabular10K<n<100K0 likes34 downloads3mo agoHugging Face02jbeiroa /resume-summarization-dataset Resume Summarization Dataset This dataset contains machine-generated summaries of 14,505 resumes using gpt-4o-mini. Each entry includes the original resume and a markdown-formatted summary divided into 5 sections. Structure Each row is a JSON object with: resume: The original resume text summary: The structured markdown summary input_tokens and output_tokens: (optional) token usage info License Some portions of this dataset are derived from public sources… See the full description on the dataset page: https://huggingface.co/datasets/jbeiroa/resume-summarization-dataset.tabularsummarization10K<n<100K1 likes22 downloads1y agoHugging Face03semeru /code-text-galeras-code-summarization-3k-dedupedtabular1K<n<10K2 likes12 downloads3y agoHugging Face04vanya-robot /russian_summarizationtabularsummarization100K<n<1M4 likes12 downloads1y agoHugging Face05divyajot5005 /plora-clean-summarization PLoRA Clean Multilingual Summarization This dataset is a cached, length-filtered training bundle for the local PLoRA notebook. It contains prompt/answer records whose full Qwen chat-formatted sequence length is at most 4096 tokens. Languages: hin_Deva (Hindi), fra_Latn (French), cmn_Hans (Chinese), urd_Arab (Urdu), eng_Latn (English), nld_Latn (Dutch), pol_Latn (Polish), snd_Arab (Sindhi), ben_Beng (Bengali), mar_Deva (Marathi) Counts: Train records: 100,000 Validation records: 8… See the full description on the dataset page: https://huggingface.co/datasets/divyajot5005/plora-clean-summarization.tabularsummarization100K<n<1M0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.