datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
taiga_corpus_subtitles
Dataset Card for Taiga Corpus - TV Series Subtitles
This dataset contains subtitles extracted from various TV series. The original data is sourced from the Taiga Corpus. It consists of line-level subtitle information with precise timing and additional metadata including series title and episode information. The dataset is designed for tasks such as subtitle alignment, translation, and dialogue analysis.
Dataset Details
Each record in the dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/Fascinat0r/taiga_corpus_subtitles.HowTo100M-subtitles-small
HowTo100M-subtitles-small
The subtitles from a subset of the HowTo100M dataset.
ERR-transcription-to-subtitlesThis dataset is created by Ilja Samoilov. In dataset is tv show subtitles from ERR and transcriptions of those shows created with TalTech ASR.
from datasets import load_dataset, load_metric
dataset = load_dataset('csv', data_files={'train': "train.tsv", \
"validation":"val.tsv", \
"test": "test.tsv"}, delimiter='\t')
Subtitles-rag-answers-r1
Subtitles-rag-answers-r1
You should mask everything except the last turn. The only part that matters to teach the model is the last turn, as you are teaching it to always output thinking, no matter what the user feeds it.
It's setup to be trained like R1:
Subtitles-rag-questions-r1
Subtitles-rag-questions-r1
You should mask everything except the last turn. The only part that matters to teach the model is the last turn, as you are teaching it to always output thinking, no matter what the user feeds it.
It's setup to be trained like R1:
open_subtitles_en_nl
Dataset Card for OpenSubtitles
Dataset Summary
This dataset is a subset from the en-nl open_subtitles dataset.
It contains only subtitles of tv shows that have a rating of at least 8.0 with at least 1000 votes.
The subtitles are also ordered and appended into buffers several lengths, with a maximum of 370 tokens
as tokenized by the 'yhavinga/ul2-base-dutch' tokenizer.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The languages… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/open_subtitles_en_nl.subtitle-summary-dataset
Subtitle Summary & Keyword Dataset (YouTube)
국내 방송사 유튜브 자막 기반 스트리밍(누적) 요약 + 유튜브 검색 키워드 데이터셋.
방송을 5분 단위로 진행하며, 매 5분마다 (직전 누적 요약 + 새 5분 자막)으로 방송 대표 한 문장 요약을 갱신하고 검색 키워드를 생성합니다.
교사 모델 Qwen3-235B-A22B-Instruct-2507-FP8 (vLLM, 자막-only). HBKenerzai/LGUplus_summary_keyword와 동일 스키마.
규모
split
회차
레코드(5분)
train
2,641
10,159
validation
294
971
합계
2,935
11,130
약 563시간 · 110개 채널 · 요약 중앙값 58자 · 도메인(유튜브 카테고리): Entertainment 2982 · Travel & Events 2081 · News &… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-dataset.survivor-subtitles-cleaned
Survivor Subtitles Dataset (cleaned)
Dataset Description
A collection of subtitles from the American reality television show "Survivor", spanning seasons 1 through 47. The dataset contains subtitle text extracted from episode broadcasts.
This dataset is a modification of the original Survivor Subtitles dataset after cleaning up and joining subtitle fragments. This dataset is a work in progress and any contributions are welcome.
Source
The subtitles were… See the full description on the dataset page: https://huggingface.co/datasets/hipml/survivor-subtitles-cleaned.subtitle-summary-lgu
Broadcast Subtitle Summary & Keyword — (LGUplus source, our pipeline)
HBKenerzai/LGUplus_summary_keyword의 방송 자막을 우리 형식으로 재구성한 뒤, **우리 파이프라인
(Qwen3-235B-A22B-Instruct-2507-FP8, vLLM, 자막-only)**으로 5분 단위 누적 요약 + 검색 키워드를
다시 생성한 데이터셋입니다. 원천 자막·메타 저작권은 원 출처(AI Hub / 방송사)에 있습니다.
규모
split
회차
레코드(5분)
train
16,609
118,608
validation
1,846
13,177
합계
18,455
131,785
약 9,025시간 분량 · 요약 길이 중앙값 63자 · 도메인: 생활정보 9528 · 정보/토크 9458 · 퀴즈/게임 8123 · 푸드/요리 7221 · 토론/대담… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-lgu.youtube_ua_noisy_subtitles_test
The list of all subsets in the dataset
Each subset is generated splitting videos from given particular ukrainiam YouTube channel
All subsets are in test split
"opodcast" subset is from channel "О! ПОДКАСТ"
"rozdympodcast" subset is from channel "Роздум | Подкаст"
"test" subset is just a small subset of samples
Loading a particular subset
>>> data_files = {"train": "data/<your_subset>.parquet"}
>>> data = load_dataset("Zarakun/youtube_ua_subtitles_test"… See the full description on the dataset page: https://huggingface.co/datasets/Zarakun/youtube_ua_noisy_subtitles_test.youtube_ua_subtitles_test
The list of all subsets in the dataset
Each subset is generated splitting videos from given particular ukrainiam YouTube channel
All subsets are in test split
"opodcast" subset is from channel "О! ПОДКАСТ"
"rozdympodcast" subset is from channel "Роздум | Подкаст"
"test" subset is just a small subset of samples
Loading a particular subset
>>> data_files = {"train": "data/<your_subset>.parquet"}
>>> data = load_dataset("Zarakun/youtube_ua_subtitles_test"… See the full description on the dataset page: https://huggingface.co/datasets/Zarakun/youtube_ua_subtitles_test.survivor-subtitles
Survivor Subtitles Dataset
Dataset Description
A collection of subtitles from the American reality television show "Survivor", spanning seasons 1 through 47. The dataset contains subtitle text extracted from episode broadcasts.
Source
The subtitles were obtained from OpenSubtitles.com.
Dataset Details
Coverage:
Seasons: 1-47
Episodes per season: ~13-14
Total episodes: ~600
Format:
Text files containing timestamped subtitle data
Character… See the full description on the dataset page: https://huggingface.co/datasets/hipml/survivor-subtitles.el-mal-el-halal-podcast-subtitles
El Mal El Halal Podcast Subtitles
Dataset Summary
El Mal El Halal Podcast Subtitles is a collection of manual subtitles for 18 episodes of the El Mal El Halal podcast by Eng. Mohamed Aboulnaga, covering Arabic content. This dataset is designed for research on speech processing, translation, semantic search, and Arabic NLP.
Total episodes: 18 - untill the date of 03/08/2025
Total segments: 13 970
Total words: 166 505
Total duration: 20 h 50 m 56 s (75 057 s)
Average… See the full description on the dataset page: https://huggingface.co/datasets/hossam87/el-mal-el-halal-podcast-subtitles.subtitle-summary-testset
Subtitle Summary & Keyword — Test set (YouTube)
방송 유튜브 자막 기반 누적 요약 + 검색 키워드 태스크의 평가용 테스트셋입니다.
jungsanghyun/subtitle-summary-dataset의 train/validation과 겹치지 않는 별도 방송으로, 동일 파이프라인(Qwen3-235B, vLLM, 자막-only)으로 생성했습니다.
규모
split
회차
레코드(5분)
test
114
519
약 25시간 · 요약 중앙값 59자 · 도메인: Entertainment 236 · News & Politics 103 · Pets & Animals 87 · Travel & Events 34 · Education 31 · Music 28
스키마
program_name, last_summary, 5min_script(입력) →… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-testset.Subtitles-rag-questions-r1-split
