datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
long-context-text-summarization-alpaca-formattext-summarizationarabic_summarization_texttext-summarization-logsro-text-summarizationnlp-summarization-image-text-mini
NLP Summarization Image Text Data Notes
Dataset summary
This data card accompanies a lightweight NLP Summarization loader for Image Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/Anandrade89/nlp-summarization-image-text-mini.nlp-summarization-image-text-clean
NLP Summarization Image Text Data Notes
Dataset summary
A documented NLP Summarization data-preparation workflow for Image Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the… See the full description on the dataset page: https://huggingface.co/datasets/HinataSato/nlp-summarization-image-text-clean.nlp-summarization-video-text-v2
NLP Summarization Video Text Data Notes
Dataset summary
This data card accompanies a lightweight NLP Summarization loader for Video Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/daniel-ramos/nlp-summarization-video-text-v2.nlp-summarization-audio-text
NLP Summarization Audio Text Data Notes
Dataset summary
This data card accompanies a lightweight NLP Summarization loader for Audio Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/adva-itpil/nlp-summarization-audio-text.summarization-yahoo-stock-finance-article-textThis is a summarization in format of key bullet points of various articles on financial stock related news from finance yahoo website.
The summarization model that was used here is llama3.3-70B
The dataset has a symbol, which is a stock that news article is related to.
dataset_131467849_nlp_summarization_video_text
dataset_131467849_nlp_summarization_video_text.py
Dataset Summary
A nlp summarization dataset with video text modality, stored in huggingface format.
Preprocessing & Augmentation
Preprocessing: minimal
Augmentation: randaugment
Splits & Sampling
Split strategy: leave one out
Sampling: stratified
Quality & Labeling
Quality filtering: moderate
Labeling: pseudo label
Files… See the full description on the dataset page: https://huggingface.co/datasets/emakarovwell/dataset_131467849_nlp_summarization_video_text.text_summarizationLegal_Text_Summarization-llama2dataset_131500778_nlp_summarization_video_text
dataset_131500778_nlp_summarization_video_text.py
Dataset Summary
A nlp summarization dataset with video text modality, stored in hdf5 format.
Preprocessing & Augmentation
Preprocessing: domain specific
Augmentation: mixup cutmix
Splits & Sampling
Split strategy: random 90 10
Sampling: contrastive
Quality & Labeling
Quality filtering: lenient
Labeling: self training
Files… See the full description on the dataset page: https://huggingface.co/datasets/rodrigomarques89/dataset_131500778_nlp_summarization_video_text.text_summarization_dataset7
Dataset Card for "text_summarization_dataset7"
More Information needed
text_summarization_dataset6
Dataset Card for "text_summarization_dataset6"
More Information needed
AMI-Corpus-Text-SummarizationDecoding-Text-Summarization-Most-Frequent-Words-and-Medical-Text-Detectionsummarization-yahoo-stock-finance-article-textThis is a summarization in format of key bullet points of various articles on financial stock related news from finance yahoo website.
The summarization model that was used here is llama3.3-70B
The dataset has a symbol, which is a stock that news article is related to.
Kaggle_CNN_Text_Summarizationtext_summarization_dataset9
Dataset Card for "text_summarization_dataset9"
More Information needed
Egyptian-text-summarization
Egyptian Arabic Text Summarization Dataset
Dataset Description
This dataset contains text-summary pairs in Egyptian Arabic designed for training and evaluating text summarization models.
Key Features
Language: Egyptian Arabic (العامية المصرية)
Task: Text Summarization
Format: Text-summary pairs with topic categorization
Content: Diverse topics with natural Egyptian Arabic usage
Dataset Structure
Data Fields
text: Original text content… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/Egyptian-text-summarization.dataset_130942876_nlp_summarization_image_text
dataset_130942876_nlp_summarization_image_text.py
Dataset Summary
A nlp summarization dataset with image text modality, stored in csv format.
Preprocessing & Augmentation
Preprocessing: curriculum
Augmentation: randaugment
Splits & Sampling
Split strategy: random 90 10
Sampling: weighted
Quality & Labeling
Quality filtering: adaptive
Labeling: semi auto
Files… See the full description on the dataset page: https://huggingface.co/datasets/jktanggraini/dataset_130942876_nlp_summarization_image_text.text_summarization_dataset1
Dataset Card for "text_summarization_dataset1"
More Information needed
synthetic-text-summarization-dataset-v1
Tanaos Text Summarization Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate text summarization systems — models that generate a concise, abstractive summary of a longer input text. It can be used to build summarization models for various applications, such as news summarization, document condensation, and content digestion.
Our flagship text summarization model… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-text-summarization-dataset-v1.text_summarization_dataset2
Dataset Card for "text_summarization_dataset2"
More Information needed
text_summarization_dataset3
Dataset Card for "text_summarization_dataset3"
More Information needed
text_summarization_dataset4
Dataset Card for "text_summarization_dataset4"
More Information needed
text_summarization_dataset5
Dataset Card for "text_summarization_dataset5"
More Information needed
text-summarization
