datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlsum-subset
Dataset Card for "XL-Sum"
Dataset Summary
We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.TSAR2025_SharedTask_RCTS_Test-Data
Citation
@inproceedings{alva-manchego-etal-2025-findings,
title = "Findings of the {TSAR} 2025 Shared Task on Readability-Controlled Text Simplification",
author = "Alva-Manchego, Fernando and Stodden, Regina and Imperial, Joseph Marvin and Barayan, Abdullah and North, Kai and Tayyar Madabushi, Harish",
editor = "Shardlow, Matthew and Alva-Manchego, Fernando and North, Kai and Stodden, Regina and Saggion, Horacio and Khallaf, Nouran and Hayakawa, Akio"… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/TSAR2025_SharedTask_RCTS_Test-Data.telugu-summarization-generation
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/telugu-summarization-generation.TSAR2025_SharedTask_RCTS_Trial-Data
Citation
@inproceedings{alva-manchego-etal-2025-findings,
title = "Findings of the {TSAR} 2025 Shared Task on Readability-Controlled Text Simplification",
author = "Alva-Manchego, Fernando and Stodden, Regina and Imperial, Joseph Marvin and Barayan, Abdullah and North, Kai and Tayyar Madabushi, Harish",
editor = "Shardlow, Matthew and Alva-Manchego, Fernando and North, Kai and Stodden, Regina and Saggion, Horacio and Khallaf, Nouran and Hayakawa, Akio"… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/TSAR2025_SharedTask_RCTS_Trial-Data.code_instruct_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
